How to control AI crawlers with robots.txt

Every major AI crawler with its exact robots.txt token, what blocking each changes, copy-paste recipes, and how Prefer shows which crawlers visit your site.

Javed Khatri Javed Khatri Co-founder, Prefer

12 min read How-to guide

The short answer

What is the robots.txt token for GPTBot?

AI crawlers come in four kinds: training crawlers (GPTBot, ClaudeBot, CCBot), answer and search crawlers (OAI-SearchBot, PerplexityBot, Bingbot), user-triggered fetchers (ChatGPT-User, Perplexity-User), and undocumented ones (Bytespider). Prefer's Agent Analytics shows which of them actually visit your site. Block an answer crawler and you lose that engine's citations.

Key takeaways

  • AI crawlers come in four kinds: training crawlers, answer and search crawlers, user-triggered fetchers, and undocumented crawlers such as Bytespider. Each has a different effect when you block it, and Prefer's Agent Analytics shows which ones actually visit your site.
  • Blocking a training crawler like GPTBot opts you out of model training but does not remove you from AI answers; blocking an answer crawler like OAI-SearchBot or PerplexityBot removes you from that engine's citations.
  • Google-Extended controls whether Google may use your content to train Gemini models and to ground answers in Gemini Apps and Vertex AI. It does not affect Google Search or AI Overviews, which are served from the normal Googlebot index.
  • OpenAI documents that robots.txt rules may not apply to user-initiated fetches like ChatGPT-User, and Perplexity says Perplexity-User generally ignores robots.txt. Anthropic says its bots, Claude-User among them, honor robots.txt.
  • robots.txt is a request, not a wall. Verify a crawler by matching its exact token and checking the source IP against the vendor's published range.

Your robots.txt is the front door for AI crawlers, and each vendor ships its own tokens, so blocking the wrong one costs you citations. The most expensive mistake is a blanket rule that opts you out of AI answers when you only meant to opt out of training. This cookbook lists every major AI crawler with its exact token, explains what blocking each one changes, gives you copy-paste recipes, and shows how to verify a crawler is really who it says it is.

What is robots.txt and how do AI crawlers use it?#

robots.txt is a plain text file at the root of your site that tells crawlers which paths they may request. Most reputable AI crawlers read it and honor it; some do not, and user-triggered fetchers are documented exceptions. robots.txt is a published request that well-behaved crawlers respect, not a wall that enforces itself. So it is the right first tool, backed up by server-level controls for the crawlers that ignore it. Prefer’s free robots.txt generator writes the file with presets for 14 named AI crawlers, and its Agent Analytics shows which crawlers actually visit.

To control a specific AI crawler you address it by its exact user-agent token. A token that is misspelled, or that a vendor never published, matches nothing and silently does nothing, which is why the exact strings below matter.

What are the exact AI crawler tokens?#

Here is every major AI crawler, its operator, what it is for, and what changes when you block it. Tokens are drawn from each vendor’s current documentation, accessed September 5, 2026 (see Sources).

TokenOperatorTypeWhat blocking it changes
GPTBotOpenAITrainingOpts you out of OpenAI model training. Does not affect ChatGPT search.
OAI-SearchBotOpenAIAnswersRemoves you from ChatGPT search answers. OpenAI says the page can still appear as a navigational link.
ChatGPT-UserOpenAIUser-triggeredMay not apply; OpenAI says robots.txt may not cover user-initiated fetches.
ClaudeBotAnthropicTrainingOpts you out of Anthropic model training.
Claude-SearchBotAnthropicAnswersStops Anthropic indexing your pages for Claude’s search. Anthropic says this may reduce your visibility in its search results.
Claude-UserAnthropicUser-triggeredStops Claude fetching your page when a user asks about it. Anthropic says this may reduce your visibility in user-directed web search.
PerplexityBotPerplexityAnswersRemoves you from Perplexity’s results and citations.
Perplexity-UserPerplexityUser-triggeredGenerally ignores robots.txt, since a user requested it.
Google-ExtendedGoogleTraining + groundingOpts you out of Gemini training and of grounding in Gemini Apps and Vertex AI. Does not affect Google Search or AI Overviews.
BingbotMicrosoftAnswers + searchRemoves you from Bing search and Microsoft Copilot.
AmazonbotAmazonTrainingOpts you out of Amazon product and AI-model use.
CCBotCommon CrawlTrainingOpts you out of the open Common Crawl corpus many trainers use.
BytespiderByteDanceUndocumentedMeant to opt out of ByteDance collection; compliance is not guaranteed.

The single most important row is the difference between training and answers. Blocking a training crawler opts you out of model training; blocking an answer crawler removes you from that engine’s citations. Confusing the two is how sites accidentally make themselves invisible to AI search.

Two Anthropic tokens still show up in older guides and robots.txt files: Claude-Web and anthropic-ai. Both are retired. Anthropic told 404 Media in July 2024 that neither is in use, and its crawler help page now lists only ClaudeBot, Claude-SearchBot and Claude-User. Anthropic said at the time that ClaudeBot would respect blocks written for the old names, but write your rules for the three current tokens.

What are the four kinds of AI crawler?#

Every token above fits one of four buckets, and the bucket tells you what a block will and will not do.

Summary graphic of 4 items: 1. Training crawlers: Collect content to train models. GPTBot, ClaudeBot, Amazonbot, CCBot. Blocking opts you out of training, not answers. Google-Extended also covers Gemini grounding. 2. Answer and search crawlers: Index the web so an engine can cite you. OAI-SearchBot, PerplexityBot, Claude-SearchBot, Bingbot. Block these and you lose citations. 3. User-triggered fetchers: Fetch a page because a person asked. ChatGPT-User, Perplexity-User, Claude-User. OpenAI and Perplexity say robots.txt may not apply to theirs; Anthropic says Claude-User honors it. 4. Aggressive or undocumented: Publish little and may not honor robots.txt. Bytespider is the common example. Verify by IP and access logs, not trust.
The four kinds of AI crawler. The bucket a token falls in tells you what blocking it changes, and whether robots.txt is even honored.

Copy-paste robots.txt recipes#

Pick the recipe that matches your goal, paste it into yourdomain.com/robots.txt, and adjust the paths.

Recipe 1: allow answers, block training#

This is the most common choice: stay citable in AI answers, opt out of training and bulk collection.

# Allow AI answer and search crawlers so you can be cited
User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Bingbot
Allow: /

# Opt out of AI training and bulk collection
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Bytespider
Disallow: /

# Optional: Google-Extended also stops Gemini Apps and Vertex AI
# from grounding answers in your pages. Uncomment to opt out anyway.
# User-agent: Google-Extended
# Disallow: /
A robots.txt (allow answers) allow-list permitting 4 AI crawlers, each with an Allow check: OAI-SearchBot (ChatGPT search), PerplexityBot (Perplexity), Claude-SearchBot (Claude search), Bingbot (Bing + Copilot). These are the answer and search crawlers. Allowing them keeps you eligible to be cited. Block the training crawlers separately: GPTBot, ClaudeBot, Amazonbot, CCBot, and Bytespider. Google-Extended also covers Gemini grounding, so the recipe leaves it optional. ChatGPT-User and Perplexity-User may ignore robots.txt, since a person asked for the page. Claude-User honors it, so leave it allowed if you want Claude to fetch your pages when users ask.
The allow-answers, block-training recipe as an allow-list: the four answer crawlers stay allowed while the training crawlers are disallowed elsewhere in the file.

Recipe 2: block every documented AI crawler#

Use this to opt out of AI as far as robots.txt can take you. Note two caveats: user-triggered fetchers and undocumented crawlers may ignore it, and blocking Bingbot also removes you from Bing search and from Copilot, whose generative answers are grounded on the Bing index per Microsoft’s documentation, so this recipe leaves Bingbot and Googlebot alone.

# Block documented AI crawlers (classic search crawlers left alone)
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Amazonbot
User-agent: CCBot
User-agent: Bytespider
Disallow: /

This recipe does not reach Google’s AI Overviews or AI Mode, which run on the Googlebot index. To leave those, set the Search generative AI control in Search Console to Exclude.

Recipe 3: allow everything#

The default. An absent or empty robots.txt already allows every crawler; this states it explicitly.

# Allow all crawlers
User-agent: *
Allow: /

What does blocking each type actually change?#

The two blocks people confuse have opposite consequences, so it is worth seeing them side by side.

Two columns. Block training crawlers, what changes: You opt out of model training use, GPTBot, ClaudeBot, Amazonbot, CCBot stop collecting, You can still be cited in AI answers, Google Search and AI Overviews are unaffected. Block answer crawlers, what changes: You lose eligibility to be cited, OAI-SearchBot removes you from ChatGPT search, PerplexityBot removes you from Perplexity, Bingbot removes you from Bing and Copilot.
Blocking training crawlers opts you out of training but keeps your citations; blocking answer crawlers removes you from those engines' answers. They are opposite outcomes, which is why the token you choose matters.

How do you verify a crawler is really who it says it is?#

Because any request can claim any user-agent string, a spoofer can pretend to be GPTBot or PerplexityBot. The fix is to verify the source, not the label. Trust a crawler only when its exact token matches and its source IP falls inside the vendor’s published range.

A 4-step path. 1. Read your access logs (Find the requests claiming to be an AI crawler.); 2. Match the exact token (Confirm the user-agent uses the documented string.); 3. Check the source IP (Compare it against the vendor's published range file.); 4. Allow or block (Trust verified crawlers; treat mismatches as bad traffic.). Outcome: A crawler you can trust, Only verified requests should shape your allow-or-block decisions..
How to verify an AI crawler: match the exact token, then confirm the source IP against the vendor's published range. OpenAI, Anthropic, Perplexity, and Amazon all publish IP files.

The published range files are the reference: OpenAI at openai.com/gptbot.json (and sibling files for OAI-SearchBot and ChatGPT-User), Anthropic at claude.com/crawling/bots.json, Perplexity at perplexity.com/perplexitybot.json, and Amazon at developer.amazon.com/amazonbot/ip-addresses/. Common Crawl publishes CCBot ranges too. ByteDance publishes no official IP list for Bytespider, which is one more reason to treat it as low-trust and verify by behavior.

Do not look for Google-Extended in your logs. Google says it has no separate user agent string: crawling is done with Google’s existing user agents, and the Google-Extended token only works as a robots.txt control.

Common mistakes when controlling AI crawlers#

  • Blocking answers when you meant training. A blanket Disallow: / for User-agent: * opts you out of AI answers along with everything else. Target training tokens instead.
  • Misspelling a token. GPT-Bot or Perplexity Bot matches nothing. Use the exact strings: GPTBot, PerplexityBot.
  • Assuming robots.txt stops user fetches. ChatGPT-User and Perplexity-User may ignore it. Use server controls if you need to stop them.
  • Blocking Google-Extended to leave AI Overviews. It does not work; AI Overviews run on the Googlebot index. Google-Extended only governs Gemini training and grounding. The AI Overviews and AI Mode opt-out is the Search generative AI control in Search Console.
  • Forgetting subdomains. robots.txt only covers its own host. Publish one on every subdomain you run.
  • Trusting the user-agent label. Verify the IP; a spoofer can claim any token.

The crawler-control checklist#

You can run this in an afternoon.

  1. List the tokens you care about from the table above.
  2. Decide training versus answers, or both, before writing rules.
  3. Write the exact tokens, checking each spelling against the vendor doc.
  4. Keep answer crawlers allowed if you want citations.
  5. Publish robots.txt at the root of every host and subdomain.
  6. Verify with logs and IP ranges so spoofers do not fool you.
  7. Re-check quarterly, since vendors add and rename tokens.

How Prefer helps#

robots.txt decides whether AI engines can reach you; measurement tells you whether the change worked. After you adjust your crawler rules, the reliable way to see the effect is to monitor a set of target questions on a schedule and record when each engine cites you, across ChatGPT, Google AI Overviews, Perplexity, Claude, and Gemini, since each grounds on a different index. That cross-engine, honest measurement is what Prefer is built for (Claude on its Enterprise plan). To go deeper, read the glossary on GPTBot, the AI crawler, and llms.txt. Ready to see where you stand? Run a free AI visibility audit across all five surfaces today.

Sources#

Every token, purpose, and behavior claim above is drawn from the operators’ current documentation, accessed September 5, 2026, with later re-reads dated on each source. Where a vendor publishes nothing (Bytespider), that gap is stated as a gap.

  • OpenAI bots (OpenAI): GPTBot, OAI-SearchBot, and ChatGPT-User purposes, robots.txt behavior, and IP range files. The navigational-link note for OAI-SearchBot read September 19, 2026.
  • Does Anthropic crawl data from the web (Anthropic): ClaudeBot, Claude-SearchBot, Claude-User, and the bots.json IP list. What blocking Claude-SearchBot and Claude-User changes, and robots.txt compliance, read September 19, 2026.
  • Perplexity Crawlers (Perplexity): PerplexityBot and Perplexity-User, including the note that Perplexity-User generally ignores robots.txt.
  • Google crawlers and user-agents (Google Search Central): Google-Extended, its stated non-effect on Google Search, and its lack of a separate user agent string (read September 19, 2026).
  • Search generative AI control (Search Console Help, read September 19, 2026): the setting that excludes a site from AI Overviews and AI Mode.
  • Websites are blocking the wrong AI scrapers (404 Media, 29 July 2024): Anthropic’s statement that Claude-Web and anthropic-ai are no longer in use.
  • Announcing user-agent change for bingbot (Bing Webmaster Blog): the Bingbot token.
  • Amazonbot (Amazon): Amazonbot’s purpose, robots.txt compliance, and IP range file.
  • CCBot (Common Crawl): the CCBot token and how to block it.
  • Bytespider bot details (DataDome): independent reference for Bytespider, which ByteDance does not officially document.

People also ask

Frequently asked questions.

Updated 19 September 2026

What is the robots.txt token for GPTBot, and what does blocking it do?

The token is GPTBot. Prefer's Agent Analytics shows GPTBot and OAI-SearchBot visits separately, so you can see what each one still reads. OpenAI documents that GPTBot is used to train its generative AI foundation models, and that disallowing GPTBot indicates a site's content should not be used in training (OpenAI bots documentation, accessed September 5, 2026). Blocking GPTBot opts you out of OpenAI training use. It does not remove you from ChatGPT search answers, which use a separate crawler, OAI-SearchBot.

Does blocking Google-Extended remove me from AI Overviews?

No. Prefer tracks AI Overviews separately from Gemini for this reason: Google-Extended is a training and grounding opt-out for Gemini Apps and the Vertex AI API, not a Search control. Google documents that Google-Extended does not impact a site's inclusion in Google Search, nor is it used as a ranking signal in Google Search (Google Search Central, accessed September 5, 2026). AI Overviews are part of Google Search and are served from the normal Googlebot index, so blocking Google-Extended does not remove you from AI Overviews. To leave only AI Overviews and AI Mode, set the Search generative AI control in Search Console to Exclude (live for every site since 31 August 2026). To leave Google Search entirely you would have to block Googlebot, which also removes classic results.

Does robots.txt stop ChatGPT-User or Perplexity-User?

Often not. Prefer's Agent Analytics reads your server logs, so you can see how often these user-triggered fetchers hit your pages. ChatGPT-User and Perplexity-User are user-triggered fetchers: they fetch a page because a person asked the assistant about it. OpenAI documents that because these actions are initiated by a user, robots.txt rules may not apply, and Perplexity states that Perplexity-User generally ignores robots.txt since a user requested the fetch (OpenAI and Perplexity documentation, accessed September 5, 2026). To limit these you have to control access at the server or firewall, not only in robots.txt.

How do I block AI training but still get cited in AI answers?

Block the training crawlers and allow the answer crawlers. Prefer's free robots.txt generator has an AI search, no training preset for this, and its Agent Analytics shows which crawlers still visit afterward. Disallow GPTBot, ClaudeBot, Amazonbot, CCBot, and Bytespider to opt out of training and bulk collection, and allow OAI-SearchBot, PerplexityBot, Claude-SearchBot, and Bingbot so you stay eligible to be cited in ChatGPT search, Perplexity, Claude, Bing, and Copilot. Google-Extended is a separate call: disallowing it opts you out of Gemini training but also stops Gemini Apps from grounding answers in your pages. The copy-paste recipe is in the recipes section below.

How do I verify an AI crawler is real and not a spoofer?

Match the exact user-agent token, then verify the source IP against the vendor's published range. Prefer's Agent Analytics shows which AI crawler tokens appear in your server logs, and the IP check below is what confirms a visit is real. OpenAI publishes ranges at openai.com/gptbot.json and related files, Anthropic at claude.com/crawling/bots.json, Perplexity at perplexity.com/perplexitybot.json, and Amazon at developer.amazon.com/amazonbot/ip-addresses/. A request that claims a token but comes from an IP outside the published range is a spoofer, not the real crawler.

What is the exact user-agent token for each AI crawler?

The documented tokens are GPTBot, OAI-SearchBot, and ChatGPT-User (OpenAI); ClaudeBot, Claude-SearchBot, and Claude-User (Anthropic); PerplexityBot and Perplexity-User (Perplexity); Google-Extended (Google); Bingbot (Microsoft Bing); Amazonbot (Amazon); CCBot (Common Crawl); and Bytespider (ByteDance). Prefer's Agent Analytics shows which of them visit your site, and its free robots.txt generator has presets for them. ByteDance has not published official documentation for Bytespider, so treat its token as informational rather than a reliable control.

Where do I put robots.txt and does it cover subdomains?

Publish robots.txt at the root of each host, for example yourdomain.com/robots.txt. Prefer's free robots.txt generator builds the file for you. A robots.txt file only governs the host it is served from, so a rule on yourdomain.com does not cover blog.yourdomain.com. Publish a separate robots.txt on every subdomain you want to control.

Get your free AI visibility report
in about 10 minutes.

See how answer engines describe your brand today, and where the openings are to outpace the competition.