Free tool

AI Crawler Access Checker can the bots behind AI answers reach you?

Paste your robots.txt and see, crawler by crawler, whether the bots that feed ChatGPT, Claude, Perplexity, Gemini and Google AI Overviews can read your pages. Fourteen named crawlers, one clear verdict each. Runs in your browser, no signup.

4/ 5 engines
GPClPxGoCo
  • ChatGPT is blocked: GPTBot is disallowed.
  • GPTBot is blocked but OAI-SearchBot is open, so ChatGPT search may still reach you while training cannot.
  • Google-Extended is disallowed: you have opted out of Gemini training. This does not affect Search or AI Overviews.
OpenAI
GPTBotTraining + answersBlockedDisallow: /
OAI-SearchBotSearch indexAllowedAllow: /, via *
ChatGPT-UserUser-triggeredAllowedAllow: /, via *
Anthropic
ClaudeBotTraining + answersAllowedAllow: /, via *
Claude-SearchBotSearch indexAllowedAllow: /, via *
Claude-UserUser-triggeredAllowedAllow: /, via *
Perplexity
PerplexityBotSearch indexAllowedAllow: /, via *
Perplexity-UserUser-triggeredAllowedAllow: /, via *
Google
GooglebotIndex + AI OverviewsAllowedAllow: /, via *
Gemini
Google-ExtendedTraining controlBlockedDisallow: /
Microsoft Copilot
BingbotIndex + CopilotAllowedAllow: /, via *
Apple
Applebot-ExtendedTraining controlAllowedAllow: /, via *
Common Crawl
CCBotOpen datasetBlockedDisallow: /
ByteDance
BytespiderTrainingAllowedAllow: /, via *
Access is the floor, not a promise of citation.

Free. Runs in your browser, so nothing you paste is uploaded. Once crawlers can reach you, check whether engines actually cite you with the AI Visibility Checker.

The short answer

Which AI crawlers you should allow.

Which AI crawlers should I allow in robots.txt?

To be found in AI answers, allow the crawlers that build them: GPTBot and OAI-SearchBot for ChatGPT, ClaudeBot for Claude, PerplexityBot for Perplexity, and Googlebot for Google AI Overviews, which use the regular Google index rather than a separate AI bot. Blocking any of these removes you from that engine's answers. The training control tokens, Google-Extended and Applebot-Extended, are a separate choice: disallowing them opts you out of model training without affecting search or AI Overviews. This checker reads a robots.txt you paste and returns a plain allowed or blocked verdict for fourteen named crawlers, plus the exact rule behind each one. It is a guide, not a guarantee: a crawler still has to choose to visit.

checked by name14

crawlers

ChatGPT to Gemini5

engines

runs in your browser0

signup

no domain fetch1

paste

What it checks

Fourteen crawlers, grouped by the engine they feed.

Each is matched to the most specific rule in your file, with Allow beating Disallow on an equal-length path, the way the crawlers themselves read it.

01
OpenAI, three bots

GPTBot for training and answers, OAI-SearchBot for ChatGPT search, and ChatGPT-User for live reads. They can be allowed or blocked independently.

02
Anthropic, three bots

ClaudeBot for the model, Claude-SearchBot for search, and Claude-User for live reads. A blanket rule catches all three.

03
Perplexity, two bots

PerplexityBot builds the index; Perplexity-User fetches live. Blocking the first removes you from Perplexity answers.

04
Google and Gemini, the important nuance

Googlebot powers both Search and AI Overviews, so blocking it costs you both. Google-Extended is a separate opt-out for Gemini training only.

05
Microsoft Copilot

Bingbot powers Bing and Copilot, and its results feed ChatGPT search, so it reaches further than its name suggests.

06
Apple Intelligence

Applebot-Extended is a training control token, not a crawler. Disallowing it opts you out of training without touching Siri indexing.

07
Open training datasets

CCBot (Common Crawl) and Bytespider (ByteDance) feed models across the field. Blocking them is a reach-versus-control trade-off.

08
Group resolution

A crawler follows the most specific user-agent group that names it, and only falls back to the wildcard group when nothing names it.

09
Longest-match rules

Within a group, the rule with the longest matching path wins, and Allow beats Disallow when two paths tie, wildcards and end-anchors included.

10
A specific path

Test the whole site or one path. A page blocked by a deep Disallow can be invisible even when the homepage is open.

How to use it

Paste robots.txt, read the verdicts.

No account, no crawl. Paste the file at yourdomain.com/robots.txt and act on the crawlers you did not mean to block.

01
Paste your robots.txt
Open yourdomain.com/robots.txt, copy all of it, and paste it in. The checker reads what you paste; it does not fetch your domain.
02
Set the path
Leave it as a slash to test the whole site, or enter a specific path to see whether one section is reachable.
03
Read the verdicts
Each of fourteen crawlers is marked allowed or blocked, grouped by engine, with the exact rule that decided it.
04
Fix the reds
Remove the Disallow lines blocking crawlers you want, republish, and re-check. Then confirm the answers follow.

Allowing a crawler lets it read you; it does not force it to visit or to cite you. Access is the floor, not the finish.

What good looks like

Open to the answer engines, deliberate about training.

A healthy setup is not open to everything. It is a clear, intentional choice per crawler.

The answer crawlers are allowed

GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Googlebot can all reach your key pages, so every major engine can read you.

Googlebot is never blocked by accident

Because AI Overviews ride the regular index, a stray Disallow on Googlebot quietly removes you from both Search and AI Overviews.

Training opt-outs are intentional

If Google-Extended or Applebot-Extended is disallowed, it is a decision you made, not a default you inherited, and you know it does not affect answers.

No crawler is blocked by a broad wildcard

A blanket Disallow under User-agent star catches AI crawlers too. Healthy files scope blocks to real paths, not everything.

Your important pages pass on their own path

The pages you want cited are reachable at their own URL, not just the homepage, with no deep Disallow hiding them.

Questions

AI crawlers and robots.txt.

Which AI crawlers should I allow?

Allow the crawlers that build answers: GPTBot and OAI-SearchBot for ChatGPT, ClaudeBot for Claude, PerplexityBot for Perplexity, and Googlebot for Google AI Overviews. Blocking any of these removes you from that engine. The training control tokens, Google-Extended and Applebot-Extended, are a separate opt-out choice that does not affect search or answers.

Does blocking GPTBot remove me from ChatGPT?

It removes you from GPTBot's crawl, which feeds OpenAI's model and answers. ChatGPT can still surface you through its search index if OAI-SearchBot and Bingbot can reach you, but you have given up the largest path. If you want to be cited in ChatGPT, allow GPTBot.

What is the difference between Googlebot and Google-Extended?

Googlebot is the crawler behind Google Search, and AI Overviews are built from that same index, so blocking Googlebot costs you both. Google-Extended is not a crawler at all. It is a control token that only decides whether your content trains Gemini and Vertex models. Disallowing it does not affect Search or AI Overviews.

Will allowing AI crawlers hurt my normal SEO?

No. Allowing GPTBot, ClaudeBot or PerplexityBot has no effect on your Google rankings; they are separate crawlers. The one that matters for both is Googlebot, which you should keep allowed for Search, AI Overviews and, indirectly, ChatGPT search.

Why can't I just enter my domain?

A browser cannot read another site's robots.txt for security reasons, and this tool runs entirely in your browser with nothing sent to a server. So you paste the file in. That also lets you test a draft before you publish it.

Is the checker free?

Yes. It is free, needs no signup, and runs entirely in your browser, so nothing you paste is uploaded. Access is only the first gate, though. Prefer's paid product is the work of turning a reachable page into an actual citation.

Get your free AI visibility report
in about 10 minutes.

See how answer engines describe your brand today, and where the openings are to outpace the competition.