Read your robots.txt for the three OpenAI crawlers#
The fastest check is Prefer’s free Crawler Access Checker; to do it by hand, open yourdomain.com/robots.txt and search for the OpenAI user agents. ChatGPT relies on three named crawlers, and each does a different job:
- OAI-SearchBot surfaces sites in ChatGPT search. Block it and ChatGPT’s search answers will not show your pages. This is the one that decides whether ChatGPT can cite you.
- GPTBot collects pages that may be used to train OpenAI’s models. Blocking it keeps you out of future training, not out of ChatGPT search.
- ChatGPT-User fetches a page when a ChatGPT user’s request needs it. OpenAI says robots.txt rules may not apply to these user-initiated visits.
Look for a Disallow line inside a User-agent: block that names any of these. A Disallow: / under User-agent: GPTBot blocks the whole site from GPTBot. A Disallow: /pricing/ blocks just that section. And a blanket Disallow: / under User-agent: * catches the AI crawlers too, which is the most common accidental block.
The rules that decide allowed versus blocked#
Crawlers do not read your file top to bottom; they resolve it by specificity, so a rule can be more or less binding than it looks. Two rules govern the outcome:
- Most specific group wins. A crawler follows the
User-agentgroup that names it. It only falls back to the wildcard*group when no block names it directly. So a permissiveGPTBotgroup can override a restrictive*group, and vice versa. - Longest match wins inside a group. The rule with the longest matching path decides, and when two paths tie,
AllowbeatsDisallow.
This is exactly where manual reading goes wrong, because a page can be blocked by a deep Disallow even when the homepage is open. To skip the guesswork, paste your file into Prefer’s free Crawler Access Checker; it evaluates 14 named AI crawlers and returns a plain allowed-or-blocked verdict with the exact rule behind each, and it runs entirely in your browser with no signup.
After you unblock: access is the floor, not the finish#
Allowing a crawler lets it read you, but it does not force it to visit or to cite you. Once your reds are fixed, republish and re-check, then confirm the answers follow over the next few weeks. If the crawlers can reach you but ChatGPT still does not name you, the problem has moved from access to visibility: the model can read you but is not choosing you.
That second gate is a content and reputation job, not a robots.txt one. Once access is clean, run a free AI visibility audit to see whether the engines actually name you, and where the gap is if they do not.
Sources
- OpenAI documents three crawlers you can control in robots.txt: GPTBot for model training, OAI-SearchBot for ChatGPT search results, and ChatGPT-User for user-initiated visits, where robots.txt rules may not apply. Each setting is independent of the others. OpenAI: Overview of OpenAI crawlers
- OpenAI publishes advertiser and publisher guidance on allowing its web crawlers through robots.txt and anti-bot layers, confirming these are the user agents to check for. OpenAI Help Center: Advertiser guidance for OpenAI web crawlers
- A crawler follows the most specific user-agent group that names it and falls back to the wildcard group only when nothing names it, and within a group the longest matching path wins with Allow beating Disallow on a tie. Prefer AI Crawler Access Checker
People also ask
- How do I know if ChatGPT can crawl my website?
- Which robots.txt rules block ChatGPT?
- Is GPTBot blocked on my site?