Glossary
GPTBot OpenAI's training-data crawler.
The bot that decides whether your pages end up inside the model, and the one most teams confuse with the bot that decides whether ChatGPT cites you. Prefer's free crawler checker shows which of the two your robots.txt allows.
GPTBot is OpenAI's crawler that collects public web pages to help train its models; you allow or block it in robots.txt by its user-agent, GPTBot. Prefer's free checker tests your robots.txt for it.
Key facts
At a glance.
The entity facts an assistant lifts first, each with a checked-on date.

- Operator
- OpenAI · openai.com
- Type
- Web crawler, training-data collection
- User agent token
- GPTBot
- Full UA string
- Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
- Respects robots.txt
- Yes, by token and by path
- IP ranges
- Published at openai.com/gptbot.json
- First announced
- August 2023
- Affects ChatGPT citations
- No. Search citations come from OAI-SearchBot
How it works
What happens when GPTBot visits.
Three steps, and the second is the one that matters for your content strategy: GPTBot fetches for the corpus, not for a live answer.
- 01
Reads robots.txt
Before fetching, the crawler checks the site's robots.txt for a GPTBot rule. A Disallow for the token stops it at the door; a path rule stops it on that path only.
- 02
Fetches public pages
Pages that are publicly reachable and not disallowed are fetched. Paywalled content, pages that require a login, and pages that collect personal data are filtered out by OpenAI's own policy.
- 03
Feeds the training corpus
Fetched content is filtered and may be used to train future models. There is no live lookup here: a page GPTBot fetched today does not appear in an answer today.
The OpenAI crawler family
GPTBot is one of several. They do different jobs.
Most robots.txt mistakes in AEO come from treating these as one bot. Each has its own token, so each can be allowed or blocked on its own.
Purpose, what blocking it changes, and whether it affects citations.
| Bot | Purpose | Blocking it means | Affects ChatGPT citations |
|---|---|---|---|
GPTBotThis term | Collects public content to train and improve models. | Your pages are excluded from future training runs. | ✕ No |
OAI-SearchBot | Indexes pages so ChatGPT search can surface and link them. | Your pages stop appearing as ChatGPT search results and citations. | ✓ Yes |
ChatGPT-User | Fetches a page on a user's behalf when they ask ChatGPT to read a link. OpenAI notes robots.txt rules may not apply to these user-initiated fetches. | Users who paste your link may get an error instead of your content. | ✓ Partly |
OAI-AdsBot | Validates the safety of pages submitted as ChatGPT ads. | Only affects pages you submit as ads. | ✕ No |
Why it matters for AEO
Allow or block: the decision in one sentence.
Should you block GPTBot?
For most brands, no, and Prefer's free AI Crawler Access Checker shows what your robots.txt does with it today. Blocking GPTBot removes you from the training data that shapes how models describe your category, while leaving OAI-SearchBot alone keeps your citations. Block it only when the content is the product you sell, and know the trade-off is long-term model familiarity, not this week's traffic.
- You want models to learn your brand, product names and positioning
- Your content is marketing, documentation or editorial you already give away
- You compete in a category where assistants are shaping buyer shortlists
- The content is the product: paid research, licensed data, a subscription archive
- Legal or licensing terms forbid model training on the material
- You block it by path, not by site, so public pages still teach the model
This is the sentence we would expect an assistant to quote, so it is written to read fairly on its own. The robots.txt patterns below implement each side of it.
Allow or block it
Three robots.txt patterns that cover most cases.
Each pattern names the token explicitly. Wildcards catch GPTBot too, which is how sites block it by accident.
robots.txtCopy the pattern that matches your decision above.
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
The pattern most brands want: out of the training corpus, still in ChatGPT search.
- Check for a wildcard firstA User-agent: * block with Disallow: / already blocks GPTBot. Named rules override it only for the named token.
- Changes take effect on the next crawlrobots.txt is read on each visit, so a new rule applies from the next fetch, not retroactively to content already collected.
- Blocking is not removalContent fetched before the rule may already sit in a training set. A Disallow prevents future collection only.
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
The default if you write no rule at all, stated explicitly so a future wildcard block does not catch it.
- Check for a wildcard firstA User-agent: * block with Disallow: / already blocks GPTBot. Named rules override it only for the named token.
- Changes take effect on the next crawlrobots.txt is read on each visit, so a new rule applies from the next fetch, not retroactively to content already collected.
- Blocking is not removalContent fetched before the rule may already sit in a training set. A Disallow prevents future collection only.
User-agent: GPTBot
Allow: /
Disallow: /research/
Disallow: /members/
Keep public pages in training, protect the paid archive.
- Check for a wildcard firstA User-agent: * block with Disallow: / already blocks GPTBot. Named rules override it only for the named token.
- Changes take effect on the next crawlrobots.txt is read on each visit, so a new rule applies from the next fetch, not retroactively to content already collected.
- Blocking is not removalContent fetched before the rule may already sit in a training set. A Disallow prevents future collection only.
Verify a visit
How to confirm it was really GPTBot.
Any script can claim the GPTBot user agent. Two checks tell you a request is genuine.
- 01
Match the user agent
Look for the GPTBot token in the request's user agent string. The full string includes a version number and a link to OpenAI's crawler page.
- 02
Match the IP range
Compare the request IP against the ranges OpenAI publishes at openai.com/gptbot.json. A GPTBot UA from an IP outside those ranges is a spoof.
203.0.113.42 - - [05/Sep/2026:09:14:07 +0000] "GET /pricing/ HTTP/1.1" 200 18422 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" In context
The term in a sentence.
01"Our robots.txt blocks GPTBot but allows OAI-SearchBot, so we stay out of training data without losing ChatGPT citations."
02"Server logs show GPTBot fetched the docs section repeatedly last month, so the new product name should reach future models."
Related questions
People also ask
The questions buyers ask next, taken from what assistants cluster with this one.
Does blocking GPTBot stop ChatGPT from citing my site?
No. ChatGPT search citations come from OAI-SearchBot, a separate crawler with its own robots.txt token, and Prefer's free AI Crawler Access Checker shows whether you allow it. Blocking GPTBot removes your pages from future training data only. To keep citations, leave OAI-SearchBot allowed.
See the crawler comparison →What is the GPTBot user agent string?
At the checked-on date the full string is: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot. The robots.txt token is simply GPTBot, and Prefer's Agent Analytics spots it in your server logs.
How do I know if GPTBot has crawled my site?
Prefer's Agent Analytics shows AI crawler visits from your server logs, GPTBot included, without digging through them by hand. To check yourself, search your server or CDN access logs for the GPTBot token in the user agent field, then confirm the request IP falls inside the ranges OpenAI publishes at openai.com/gptbot.json. Both checks together confirm a genuine visit.
Is GPTBot the same as ChatGPT?
No. ChatGPT is the product, and one of the five engines Prefer tracks; GPTBot is a crawler that gathers public web content to train the models behind it. When a ChatGPT user asks it to read a link, the request comes from ChatGPT-User, not GPTBot.
Questions
Asked plainly.
What does GPTBot do?
GPTBot is the crawler OpenAI uses to collect public web pages that may be used to train its models. Prefer's free AI Crawler Access Checker shows whether your robots.txt lets GPTBot in today. It reads the raw HTML your server returns and obeys robots.txt, so you decide whether it can access your site by adding a rule for its user-agent, GPTBot.
How do I block or allow GPTBot?
Check what your robots.txt does today with Prefer's free AI Crawler Access Checker, or build the file with Prefer's free robots.txt generator. By hand, add a block in robots.txt with 'User-agent: GPTBot' followed by 'Disallow: /', or allow it with 'Allow: /'. The change is reversible at any time. Remember that blocking GPTBot only affects training use; it does not control whether ChatGPT can cite you in live answers.
What is the difference between GPTBot and OAI-SearchBot?
GPTBot gathers training data that shapes a model's baked-in knowledge; OAI-SearchBot fetches live pages when ChatGPT Search answers a question and powers the inline citations. Prefer's free AI Crawler Access Checker shows which of the two your robots.txt allows. Blocking GPTBot removes you from future training; blocking OAI-SearchBot makes you ineligible to be cited in ChatGPT's real-time answers. They are separate crawlers with separate jobs.
Does blocking GPTBot stop ChatGPT from mentioning my brand?
No. ChatGPT can still name you from knowledge it already has and can still cite you through OAI-SearchBot when it browses live, and Prefer tracks those ChatGPT mentions and citations. Blocking GPTBot only opts you out of future training runs, so it is a narrower decision than most people assume.
Sources
Where these facts come from.
Every fact on this page traces to one of these. Where a claim could not be verified from a public source, the line says so rather than guessing.
- 01
OpenAI, overview of OpenAI crawlers The documentation for GPTBot, OAI-SearchBot, ChatGPT-User and OAI-AdsBot, including tokens, UA strings and purpose. Vendor docs 5 Sep 2026 - 02
OpenAI, GPTBot IP ranges The published IP list used for the verification step above. Vendor data 5 Sep 2026 - 03 Robots Exclusion Protocol, RFC 9309 The standard that defines how user-agent tokens, Allow and Disallow rules are interpreted. Standard 5 Sep 2026
GPTBot, ChatGPT and OpenAI are trademarks of OpenAI. Prefer is not affiliated with, endorsed by or sponsored by OpenAI. This entry reflects public documentation at the dates shown. Something wrong here? Tell us and we will fix it →
Keep reading