resulthack

Which AI-training opt-out signals does your site send?

There are at least five ways to tell AI companies not to train on your site, each defined by someone different. Paste your files and see which signals you send, whether each is written the way its definition says, and which AI operators say in their own documentation that they read it. Everything is read in this tab and nothing is fetched or sent.

Who says they honour which signal

Each cell comes from the operator's own documentation, quoted, with the page and the date we read it. "Not documented" means that page does not mention the signal: it is the absence of a statement, not proof the signal is ignored, and an operator may read a signal it does not write about. It also cannot tell you whether a crawler actually complies. The same table is in opt-out-signals.json.

OperatorPer-agent robots.txt tokensContent SignalsTDMRepnoai and noimageaiai.txt
OpenAI“Not documented”: not mentioned on OpenAI’s crawler page, read 25 Sep 2026Per-agent robots.txt tokensDocumented GPTBotDisallowing GPTBot indicates a site’s content should not be used in training generative AI foundation models.OpenAI documentation, read 25 Sep 2026Content SignalsNot documentedTDMRepNot documentednoai and noimageaiNot documentedai.txtNot documented
Anthropic“Not documented”: not mentioned on Anthropic’s crawler page, read 25 Sep 2026Per-agent robots.txt tokensDocumented ClaudeBotWhen a site restricts ClaudeBot access, it signals that the site's future materials should be excluded from our AI model training datasets.Anthropic documentation, read 25 Sep 2026Content SignalsNot documentedTDMRepNot documentednoai and noimageaiNot documentedai.txtNot documented
Google“Not documented”: not mentioned on Google’s crawler page, read 25 Sep 2026Per-agent robots.txt tokensDocumented Google-ExtendedGoogle-Extended is a standalone product token that web publishers can use to manage whether content Google crawls from their sites may be used for training future generations of Gemini models that power Gemini Apps and Vertex AI API for Gemini and for grounding (providing content from the Google Search index to the model at prompt time to improve factuality and relevancy) in Gemini Apps and Grounding with Google Search on Vertex AI.Google-Extended also controls grounding in Gemini Apps, not only training. Google says it does not affect inclusion in Google Search.Google documentation, read 25 Sep 2026Content SignalsNot documentedTDMRepNot documentednoai and noimageaiNot documentedai.txtNot documented
Apple“Not documented”: not mentioned on Apple’s crawler page, read 25 Sep 2026Per-agent robots.txt tokensDocumented Applebot-ExtendedWith Applebot-Extended, web publishers can choose to opt out of their website content being used to train Apple’s general purpose foundation models powering generative AI features across Apple products, including Apple Intelligence, Services, and Developer Tools.Apple documentation, read 25 Sep 2026Content SignalsNot documentedTDMRepNot documentednoai and noimageaiNot documentedai.txtNot documented
Meta“Not documented”: not mentioned on Meta’s crawler page, read 25 Sep 2026Per-agent robots.txt tokensDocumented Meta-ExternalAgentIn order to block these crawlers, add a disallow for the relevant crawler to robots.txt.Meta documentation, read 25 Sep 2026Content SignalsNot documentedTDMRepNot documentednoai and noimageaiPrefers robots.txtWe make it easy for site managers and content owners to indicate their preferences by using industry-standard practices like robots.txt rather than non-standard formats like NoAI tags.Meta documentation, read 25 Sep 2026ai.txtNot documented
Amazon“Not documented”: not mentioned on Amazon’s crawler page, read 25 Sep 2026Per-agent robots.txt tokensDocumented AmazonbotAutomated crawling from these listed user agents respects the Robots Exclusion Protocol , honoring the user-agent and the allow/disallow directives.Amazon documentation, read 25 Sep 2026Content SignalsNot documentedTDMRepNot documentednoai and noimageaiNot documentedai.txtNot documented
Mistral AI“Not documented”: not mentioned on Mistral AI’s crawler page, read 25 Sep 2026Per-agent robots.txt tokensDocumented MistralAI-TrainingMistralAI-Training crawls web content to help build datasets for training Mistral generative AI models. Webmasters can disallow this user agent in their robots.txt file.Mistral AI documentation, read 25 Sep 2026Content SignalsNot documentedTDMRepNot documentednoai and noimageaiNot documentedai.txtNot documented
Common Crawl“Not documented”: not mentioned on Common Crawl’s crawler page, read 25 Sep 2026Per-agent robots.txt tokensDocumented CCBotAdd these lines to your robots.txt file and our crawler will stop crawling your website:This stops CCBot crawling at all. Common Crawl’s own pages do not describe CCBot as an AI-training crawler; its dataset is widely used for training by others.Common Crawl documentation, read 25 Sep 2026Content SignalsNot documentedTDMRepNot documentednoai and noimageaiNot documentedai.txtNot documented

Only one kind of signal has written support from all eight: a robots.txt line naming the operator's own token. Meta is the only one of the eight that addresses noai at all, and it points to robots.txt instead. None of the eight crawler pages mentions Content-Signal, TDMRep or ai.txt.

The five signals, and who defined each

Per-agent robots.txt tokens

Where: robots.txt at the site root.

Defined by: robots.txt is RFC 9309 (IETF, 2022). Each training token is defined by the operator that reads it: GPTBot by OpenAI, Google-Extended by Google, and so on.

Specification: RFC 9309: Robots Exclusion Protocol (read 25 Sep 2026).

Content Signals (Content-Signal line in robots.txt)

Where: robots.txt, inside a User-agent group.

Defined by: Cloudflare, the Content Signals Policy, announced 24 Sep 2025. It defines search, ai-input and ai-train (yes or no); Cloudflare's robots.txt documentation (last updated 3 Aug 2026) adds a fourth field, use, with the values immediate, reference or full.

Specification: Content Signals (contentsignals.org, the policy text and examples) (read 25 Sep 2026); Cloudflare blog: Giving users choice with Cloudflare's new Content Signals Policy (24 Sep 2025) (read 25 Sep 2026); Cloudflare docs: robots.txt setting (use= field; last updated 3 Aug 2026) (read 25 Sep 2026).

TDMRep (TDM Reservation Protocol)

Where: /.well-known/tdmrep.json, the tdm-reservation and tdm-policy HTTP headers, and meta tags of the same names.

Defined by: The W3C Text and Data Mining Reservation Protocol Community Group; Final Community Group Report of 10 May 2024, editor Laurent Le Meur (EDRLab). The report itself says it is not a W3C Standard nor on the W3C Standards Track.

Specification: TDM Reservation Protocol (TDMRep), Final Community Group Report, 10 May 2024 (read 25 Sep 2026).

noai and noimageai

Where: a robots meta tag, or an X-Robots-Tag response header.

Defined by: DeviantArt, November 2022, for its own artists' pages; other sites copied it. It is not in Google's or Bing's robots meta documentation and has no published specification.

Specification: DeviantArt: UPDATE All Deviations Are Opted Out of AI Datasets (Nov 2022) (read 25 Sep 2026; the page answered our fetch with HTTP 403; the origin and date are as the page is indexed by search engines and reported by the press at the time).

ai.txt (Spawning)

Where: /ai.txt at the site root.

Defined by: Spawning, the company behind Have I Been Trained and the Do Not Train registry. It uses robots.txt syntax (User-Agent, Allow, Disallow with file-type patterns). There is no published specification beyond Spawning's generator and FAQ; Spawning says its API passes ai.txt permissions to "a growing list of AI researchers and partners" without naming them.

Specification: Spawning: ai.txt generator and FAQ (read 25 Sep 2026).

How TDMRep's three places combine

TDMRep section 6.7 sets the order: a TDM agent reads /.well-known/tdmrep.json first, then the tdm-reservation header on each response, then the meta tag in each HTML page. Each later value replaces the earlier one, and a missing value never resets it. So a site-wide file saying 1 and a page header saying 0 means 0 for that page. In tdmrep.json, the first rule whose location matches the path wins, and patterns use robots.txt syntax (* and a final $). tdm-reservation is 1 (reserved) or 0 (not reserved); the report calls any other value a protocol error, to be treated as unset.

Content-Signal, line by line

A Content-Signal line sits inside a User-agent group of robots.txt, so it applies to the agents that group names. Its value is comma-separated key=value pairs: search, ai-input and ai-train take yes or no, and use (in Cloudflare's robots.txt documentation, last updated 3 Aug 2026) takes immediate, reference or full. A path before the pairs, as in Content-Signal: /blog/ ai-train=no, scopes the line to that path, following contentsignals.org's own examples. Leaving a key out expresses no preference for that use, in the policy's words it "neither grants nor restricts permission".

The law these signals point to

TDMRep, Content Signals and Spawning's ai.txt all say they answer Article 4 of the EU's 2019 copyright directive. Its paragraph 3 reads:

“The exception or limitation provided for in paragraph 1 shall apply on condition that the use of works and other subject matter referred to in that paragraph has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online.”

Directive (EU) 2019/790, Article 4(3), EUR-Lex. Read 25 Sep 2026 from the Publications Office copy of the official text, because EUR-Lex answered our plain request with a browser challenge.

This page is not legal advice. It does not say whether any signal, or any combination, is a reservation "in an appropriate manner" under that article, in your country or anywhere else, and it does not cover the law outside the EU. The directive does not name any file or header. If it matters to you, ask a lawyer who practises in your jurisdiction.

What this checker cannot do