AI crawler list

Every AI crawler, what it does, and whether to let it in.

Prefer keeps this list of the AI bots that read your site: who runs each one, what it fetches for, and how to allow or block it. 20 crawlers from 12 operators, each fact taken from the operator's own documentation and checked on 1 October 2026.

20 crawlers · Checked 1 October 2026

The list

AI crawlers by operator and purpose

The purpose column is the one that matters for AI visibility. Prefer's Agent Analytics reads your server logs and shows which of these bots actually visit, and the free Crawler Access Checker tells you which ones your robots.txt blocks today.

CrawlerOperatorPurposeRespects robots.txtOfficial docs
Amazonbot Amazonbot Amazon Mixed Yes Docs · IPs
ClaudeBot ClaudeBot Anthropic Training Yes Docs · IPs
Claude-SearchBot Claude-SearchBot Anthropic Search Yes Docs · IPs
Claude-User Claude-User Anthropic User-triggered Yes Docs · IPs
Applebot-Extended Applebot-Extended Apple Training Yes Docs
Applebot Applebot Apple Mixed Yes Docs · IPs
Bytespider Bytespider ByteDance Undocumented Not documented None published
CCBot CCBot Common Crawl Foundation Training Yes Docs · IPs
DuckAssistBot DuckAssistBot DuckDuckGo Search Yes Docs · IPs
Googlebot Googlebot Google Search Yes Docs · IPs
Google-Extended Google-Extended Google Mixed Yes Docs
Meta-ExternalFetcher meta-externalfetcher Meta User-triggered Partly Docs
Meta-ExternalAgent meta-externalagent Meta Mixed Yes Docs
Bingbot bingbot Microsoft Mixed Yes Docs · IPs
MistralAI-User MistralAI-User Mistral AI User-triggered Yes Docs · IPs
GPTBot GPTBot OpenAI Training Yes Docs · IPs
OAI-SearchBot OAI-SearchBot OpenAI Search Yes Docs · IPs
ChatGPT-User ChatGPT-User OpenAI User-triggered Partly Docs · IPs
PerplexityBot PerplexityBot Perplexity Search Yes Docs · IPs
Perplexity-User Perplexity-User Perplexity User-triggered No Docs · IPs
Training
Collects pages that may train models. Blocking it does not, by itself, remove you from live AI answers.
Search
Builds the index an AI answer engine cites from. Blocking it can keep you out of those answers.
User-triggered
Fetches a page because a person asked the assistant to. Some operators say these fetches may not follow robots.txt.
Mixed
One crawler serving more than one job, for example search and AI features.
Undocumented
The operator publishes no documentation for this crawler.

robots.txt starter

Allow the bots that cite you, decide on the ones that train

Prefer's starting point for a brand that wants AI visibility: let the search and user-triggered fetchers in, and make a deliberate choice about training crawlers. Tokens are copied from each operator's documentation. Test the result with the free Crawler Access Checker or build your own file with the robots.txt generator.

# AI search and user-triggered fetchers: allow, so AI answers can cite you
User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: DuckAssistBot
Allow: /

User-agent: Googlebot
Allow: /

User-agent: meta-externalfetcher
Allow: /

User-agent: MistralAI-User
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

# Training crawlers: your choice. Blocking them keeps pages out of future
# training, not out of live citations. Remove the # to block.
# User-agent: ClaudeBot
# Disallow: /

# User-agent: Applebot-Extended
# Disallow: /

# User-agent: CCBot
# Disallow: /

# User-agent: GPTBot
# Disallow: /

Two tokens on this list do not crawl at all: Google-Extended and Applebot-Extended are opt-out signals read by Googlebot and Applebot. Googlebot itself controls AI Overviews and AI Mode, so blocking it removes you from Google Search as well; leave it allowed.

Get your free AI visibility report
in about 10 minutes.

See how answer engines describe your brand today, and where the openings are to outpace the competition.