AI crawler list
Every AI crawler, what it does, and whether to let it in.
Prefer keeps this list of the AI bots that read your site: who runs each one, what it fetches for, and how to allow or block it. 20 crawlers from 12 operators, each fact taken from the operator's own documentation and checked on 1 October 2026.
The list
AI crawlers by operator and purpose
The purpose column is the one that matters for AI visibility. Prefer's Agent Analytics reads your server logs and shows which of these bots actually visit, and the free Crawler Access Checker tells you which ones your robots.txt blocks today.
| Crawler | Operator | Purpose | Respects robots.txt | Official docs |
|---|---|---|---|---|
Amazonbot Amazonbot | Amazon | Mixed | Yes | Docs · IPs |
ClaudeBot ClaudeBot | Anthropic | Training | Yes | Docs · IPs |
Claude-SearchBot Claude-SearchBot | Anthropic | Search | Yes | Docs · IPs |
Claude-User Claude-User | Anthropic | User-triggered | Yes | Docs · IPs |
Applebot-Extended Applebot-Extended | Apple | Training | Yes | Docs |
Applebot Applebot | Apple | Mixed | Yes | Docs · IPs |
Bytespider Bytespider | ByteDance | Undocumented | Not documented | None published |
CCBot CCBot | Common Crawl Foundation | Training | Yes | Docs · IPs |
DuckAssistBot DuckAssistBot | DuckDuckGo | Search | Yes | Docs · IPs |
Googlebot Googlebot | Search | Yes | Docs · IPs | |
Google-Extended Google-Extended | Mixed | Yes | Docs | |
Meta-ExternalFetcher meta-externalfetcher | Meta | User-triggered | Partly | Docs |
Meta-ExternalAgent meta-externalagent | Meta | Mixed | Yes | Docs |
Bingbot bingbot | Microsoft | Mixed | Yes | Docs · IPs |
MistralAI-User MistralAI-User | Mistral AI | User-triggered | Yes | Docs · IPs |
GPTBot GPTBot | OpenAI | Training | Yes | Docs · IPs |
OAI-SearchBot OAI-SearchBot | OpenAI | Search | Yes | Docs · IPs |
ChatGPT-User ChatGPT-User | OpenAI | User-triggered | Partly | Docs · IPs |
PerplexityBot PerplexityBot | Perplexity | Search | Yes | Docs · IPs |
Perplexity-User Perplexity-User | Perplexity | User-triggered | No | Docs · IPs |
- Training
- Collects pages that may train models. Blocking it does not, by itself, remove you from live AI answers.
- Search
- Builds the index an AI answer engine cites from. Blocking it can keep you out of those answers.
- User-triggered
- Fetches a page because a person asked the assistant to. Some operators say these fetches may not follow robots.txt.
- Mixed
- One crawler serving more than one job, for example search and AI features.
- Undocumented
- The operator publishes no documentation for this crawler.
robots.txt starter
Allow the bots that cite you, decide on the ones that train
Prefer's starting point for a brand that wants AI visibility: let the search and user-triggered fetchers in, and make a deliberate choice about training crawlers. Tokens are copied from each operator's documentation. Test the result with the free Crawler Access Checker or build your own file with the robots.txt generator.
# AI search and user-triggered fetchers: allow, so AI answers can cite you
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: DuckAssistBot
Allow: /
User-agent: Googlebot
Allow: /
User-agent: meta-externalfetcher
Allow: /
User-agent: MistralAI-User
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
# Training crawlers: your choice. Blocking them keeps pages out of future
# training, not out of live citations. Remove the # to block.
# User-agent: ClaudeBot
# Disallow: /
# User-agent: Applebot-Extended
# Disallow: /
# User-agent: CCBot
# Disallow: /
# User-agent: GPTBot
# Disallow: / Two tokens on this list do not crawl at all: Google-Extended and Applebot-Extended are opt-out signals read by Googlebot and Applebot. Googlebot itself controls AI Overviews and AI Mode, so blocking it removes you from Google Search as well; leave it allowed.
Go deeper