Free tool

Robots.txt Generator with AI crawler presets

Choose which AI crawlers can read your site, list the paths you want kept private, and get a robots.txt that says exactly that. Every file is checked with the same rules as our Crawler Access Checker. Runs in your browser, no signup.

Start from a preset

Every search, user-triggered and training crawler is allowed. The simplest choice for a brand that wants to be found and remembered by AI.

OpenAI
GPTBotTraining
OAI-SearchBotSearch index
ChatGPT-UserUser-triggered
Anthropic
ClaudeBotTraining
Claude-SearchBotSearch index
Claude-UserUser-triggered
Perplexity
PerplexityBotSearch index
Perplexity-UserUser-triggered
Google
GooglebotIndex + AI Overviews
Gemini
Google-ExtendedTraining + grounding
Microsoft Copilot
BingbotIndex + Copilot
Apple
Applebot-ExtendedTraining control
Common Crawl
CCBotOpen dataset
ByteDance
BytespiderTraining
5/ 5 answer engines can reach /
ChatGPTClaudePerplexityGoogle AI OverviewsCopilot
  • All five answer engines can reach /.
robots.txt
# robots.txt

# Every other crawler
User-agent: *
Disallow: /admin/
Disallow: /cart/

# AI crawlers you allow. A crawler named in its own group ignores the
# User-agent: * group, so the private paths are repeated here.
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Googlebot
User-agent: Google-Extended
User-agent: Bingbot
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Bytespider
Disallow: /admin/
Disallow: /cart/

Sitemap: https://example.com/sitemap.xml

Free. Runs in your browser, so nothing you enter is uploaded. Already have a robots.txt? Paste it into the Crawler Access Checker for a verdict on each crawler.

The short answer

What robots.txt should say about AI crawlers.

What should robots.txt say about AI crawlers?

Robots.txt should name the AI crawlers you want in and the ones you want out, because the major operators split search from training. To appear in AI search answers, allow OAI-SearchBot for ChatGPT, Claude-SearchBot for Claude, PerplexityBot for Perplexity, Bingbot for Copilot and Googlebot, which also feeds Google AI Overviews. Training is a separate choice: GPTBot and ClaudeBot collect training data, and Applebot-Extended is an opt-out token, so you can block those and stay in search answers. Google-Extended is an opt-out token too, but it also covers grounding in Gemini Apps, so blocking it keeps Gemini Apps from grounding answers in your pages. One rule trips people up: a crawler named in its own group ignores the rules under User-agent: *, so private paths have to be repeated in that group. This generator does that for you and checks the file before you copy it.

named in the file14

AI crawlers

plus custom4

presets

per OpenAI~24 h

for OpenAI to read changes

runs in your browser0

signup

What it builds

A readable robots.txt, checked before you copy it.

The file uses the standard robots.txt rules every major crawler follows, grouped so the next person to open it can see what you decided.

01
Presets for the common choices

Open to all AI, AI search without training, block bulk datasets, or block AI while keeping Google and Bing. Start from one, then change any crawler.

02
Fourteen named crawlers

The same registry as the Crawler Access Checker, grouped by operator, each labelled with its documented role: search, user-triggered, training or training control.

03
Private paths, repeated correctly

Paths you list are blocked for every crawler, and repeated inside each named group, because a crawler with its own group ignores the rules under User-agent: *.

04
Grouped, commented output

Allowed crawlers share one group and blocked crawlers share another, each with a comment above it, instead of fourteen near-identical blocks.

05
A check before you copy

The file is parsed with the checker's rules, and you can test any path to see how many of the five answer engines can reach it.

06
Your sitemap

Add your sitemap URL and it is written at the end of the file, where crawlers look for it.

How to use it

From a preset to a published file in about two minutes.

No account and nothing to install. Publish the result at the root of your domain.

01
Pick a preset
Choose the posture closest to what you want. You can change any single crawler afterwards.
02
Adjust single crawlers
Switch any crawler between allow and block. The preset changes to Custom, so you can see you have left it.
03
Add private paths and your sitemap
List paths no crawler should read, like /admin/ or /cart/, one per line, then add your sitemap URL.
04
Check, copy, publish
Test a path, confirm the engines you want can reach it, then copy or download the file and publish it as /robots.txt at your domain root.

Robots.txt is a request, not a lock. Well-behaved crawlers follow it, and user-triggered fetchers may not. To enforce a block, use your firewall or CDN bot rules.

What good looks like

Open to AI search, deliberate about the rest.

A good robots.txt is not the longest one. It is the one where every AI crawler's access is a choice somebody made.

The search crawlers are allowed on purpose

OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot can reach the pages you want cited.

Training choices are explicit

GPTBot, ClaudeBot and Applebot-Extended each get a clear allow or block, not a default you inherited from a CMS. So does Google-Extended, which also decides whether Gemini Apps can ground answers in your pages.

Private paths hold for every group

A path you block under User-agent: * is repeated in every named group, so naming a crawler never quietly opens /admin/ to it.

Googlebot and Bingbot stay open

Blocking either removes you from Google or Bing search, and from AI Overviews or Copilot along with it.

One file per host

robots.txt applies only to the host it sits on, so blog.example.com needs its own file at blog.example.com/robots.txt.

Questions

robots.txt for AI crawlers.

How do I allow AI crawlers in robots.txt?

Name them in a User-agent group with Allow: /, or with only the Disallow lines for paths you want private. The crawlers that put you in AI search answers are OAI-SearchBot for ChatGPT, Claude-SearchBot for Claude, PerplexityBot for Perplexity, Bingbot for Copilot and Googlebot for Google AI Overviews. If nothing in your file blocks them they are already allowed, but naming them makes the intent clear.

Can I block AI training and still appear in AI search?

Yes. OpenAI says its crawler settings are independent, so you can disallow GPTBot and allow OAI-SearchBot. Do the same with ClaudeBot and Claude-SearchBot for Anthropic, and disallow Applebot-Extended to opt out of Apple Intelligence training. Disallowing Google-Extended opts you out of Gemini training but also of grounding in Gemini Apps, so weigh it separately. Perplexity says PerplexityBot is not used to train foundation models. The AI search, no training preset sets this up, and it blocks Google-Extended too.

Why are my private paths repeated for each group?

Because a crawler follows only the most specific group that names it. If GPTBot has its own group, it ignores every rule under User-agent: *, including your Disallow: /admin/. Repeating the paths inside each named group keeps them private for every crawler.

Does robots.txt stop every AI bot?

No. It is a request that well-behaved crawlers follow. OpenAI says robots.txt rules may not apply to ChatGPT-User because a person started the request, and Perplexity says Perplexity-User generally ignores robots.txt. Some scrapers ignore it entirely. To enforce a block, use your firewall, CDN bot rules or a login.

How long until AI crawlers see my changes?

OpenAI says it can take about 24 hours from a robots.txt update for its systems to adjust. Other operators do not publish a figure, so re-check the live file with the Crawler Access Checker after publishing and allow a few days before judging the effect.

Where does robots.txt go?

At the root of each host, as /robots.txt, for example https://example.com/robots.txt. Crawlers do not look for it in subfolders, and each subdomain needs its own file.

Is the robots.txt generator free?

Yes. It is free, needs no signup, and runs entirely in your browser, so nothing you enter is uploaded. Access is only the first step. Prefer's paid plans measure whether ChatGPT, Gemini, Perplexity and Google AI Overviews then cite you, and do the work to change that.

Get your free AI visibility report
in about 10 minutes.

See how answer engines describe your brand today, and where the openings are to outpace the competition.