Glossary

ClaudeBot Anthropic's training-data crawler.

  • Noun
  • Crawler
  • Owned by Anthropic

The Anthropic bot that collects web content for model training, and the one most often confused with the bots behind Claude's search. Prefer's free crawler checker shows how your robots.txt treats all three Anthropic bots.

Updated1 Oct 2026Reading time5 min
Definition32 words

ClaudeBot is Anthropic's crawler that collects public web content that could contribute to training its AI models; you allow or block it in robots.txt as ClaudeBot. Prefer's free checker tests for it.

Also writtenAnthropic crawlerAnthropic training botClaude crawler
Key facts ↓

Key facts

At a glance.

ClaudeBot at a glance

The entity facts an assistant lifts first, each checked against Anthropic's own documentation.

Checked1 Oct 2026
Operator
Anthropic · anthropic.com
Type
Web crawler, training-data collection
User agent token
ClaudeBot
Full UA string
Not documented by Anthropic
Respects robots.txt
Yes, plus the non-standard Crawl-delay extension
IP ranges
Published at claude.com/crawling/bots.json (one list for all Anthropic bots)
First announced
Not documented
Affects Claude answers
Not documented. Anthropic ties ClaudeBot to training datasets only

How it works

What happens when ClaudeBot visits.

Three steps. The third is the one to remember: what ClaudeBot collects is about training, not a live answer.

  1. 01

    Reads robots.txt

    Anthropic says its bots honour industry standard robots.txt directives and support Crawl-delay. A Disallow for ClaudeBot is the opt-out, set per subdomain.

  2. 02

    Collects public content

    Pages it is allowed to reach are collected. Anthropic says its bots respect anti-circumvention technologies and will not try to bypass CAPTCHAs.

  3. 03

    May contribute to training

    Collected content could potentially contribute to training Anthropic's generative AI models. Restricting ClaudeBot signals your future materials should be excluded from training datasets.

NoteAnthropic's wording is about future materials. A Disallow added today does not, on Anthropic's documentation, reach content collected before it.

The Anthropic crawler family

ClaudeBot is one of three. They do different jobs.

Anthropic documents three robots, each with its own robots.txt token, so each can be allowed or blocked on its own.

Anthropic's bots

Purpose in Anthropic's words, what blocking changes, and whether Anthropic says it affects search visibility.

Checked1 Oct 2026
BotPurposeBlocking it meansAffects search visibility
ClaudeBotThis term Collects web content that could potentially contribute to training Anthropic's generative AI models. Signals your future materials should be excluded from training datasets. ✕ Not documented
Claude-SearchBot Navigates the web to improve search result quality for users. Your content is not indexed for search, which may reduce your visibility in user search results. ✓ Yes, may reduce it
Claude-User Accesses websites when individuals ask Claude questions. Your content is not retrieved for user queries, which may reduce visibility for user-directed web search. ✓ Yes, may reduce it
BasisTokens and purposes as documented by Anthropic in its Help Center article on web crawling at the checked-on date. Anthropic does not publish full user agent strings for these bots.

Why it matters for AEO

Allow or block: the decision in one sentence.

Should you block ClaudeBot?

For most brands that want AI visibility, no, and Prefer's free AI Crawler Access Checker shows what your robots.txt does with it today. ClaudeBot is Anthropic's training crawler, so blocking it signals that your future pages should be excluded from the training data that shapes how Claude describes your category. Anthropic documents no effect on search visibility for ClaudeBot; that is Claude-SearchBot's job.

Allow ClaudeBot if
  • You want Anthropic's models to learn your brand, products and positioning
  • Your content is marketing, documentation or editorial you already give away
  • Buyers in your category use Claude to research and shortlist
Block ClaudeBot if
  • The content is the product: paid research, licensed data, a subscription archive
  • Licensing terms forbid model training on the material
  • You block by path, and leave Claude-SearchBot and Claude-User allowed

If crawl load is the worry rather than training, Anthropic supports Crawl-delay, which slows ClaudeBot without blocking it.

Allow or block it

Three robots.txt patterns that cover most cases.

Each names the ClaudeBot token explicitly. Anthropic asks you to add the rule on every subdomain you want covered.

robots.txt

Copy the pattern that matches your decision above.

# Opt out of training
User-agent: ClaudeBot
Disallow: /

# Search indexing for Claude users
User-agent: Claude-SearchBot
Allow: /

# Pages fetched when a Claude user asks
User-agent: Claude-User
Allow: /
Block training, keep search

Out of future training data, still reachable by the bots for Claude's search and user requests.

  • Repeat it per subdomainAnthropic says to add the rule to the robots.txt of every subdomain you want to opt out. A rule on www does not cover docs.
  • Do not rely on IP blockingAnthropic warns that blocking its IP addresses may not work reliably, because it stops ClaudeBot reading your robots.txt.
  • Check for a wildcard firstA User-agent: * group with Disallow: / already blocks ClaudeBot unless a group names ClaudeBot directly.

Verify a visit

How to confirm it was really ClaudeBot.

Anthropic publishes the robots.txt token and one IP list, but no full user agent string. The IP is the check that counts.

  1. 01

    Look for the token

    A request naming ClaudeBot in its user agent is a claim, not proof. Anthropic does not publish the full string, so do not match on a string copied from a third-party list.

  2. 02

    Match the IP list

    Compare the request IP with claude.com/crawling/bots.json. Anthropic says a crawler with a source IP on this list is coming from Anthropic. The list covers all three Anthropic bots.

  3. 03

    Contact Anthropic

    Anthropic lists [email protected] as its crawler contact and asks you to write from an email that includes the domain you are contacting it about.

In context

The term in a sentence.

01

"We block ClaudeBot on the paid archive only, and leave Claude-SearchBot allowed site-wide."

02

"ClaudeBot was hitting the docs hard, so we added Crawl-delay instead of blocking it."

Related questions

People also ask

The questions buyers ask next, taken from what assistants cluster with this one.

Does blocking ClaudeBot remove me from Claude's search results?

Anthropic does not document that, and Prefer's free AI Crawler Access Checker shows the separate rule for Claude-SearchBot. Anthropic ties ClaudeBot to training datasets only. It says disabling Claude-SearchBot is what may reduce your visibility in user search results.

See the crawler comparison →
What is the ClaudeBot user agent string?

Anthropic does not publish one, and Prefer's Agent Analytics identifies ClaudeBot from your server logs without you needing it. Anthropic documents only the robots.txt token, ClaudeBot, and an IP list at claude.com/crawling/bots.json for checking that a visit is genuine.

How do I know if ClaudeBot has crawled my site?

Prefer's Agent Analytics shows AI crawler visits from your server logs, ClaudeBot included. To check by hand, search your access logs for ClaudeBot, then confirm the request IP is on the list Anthropic publishes at claude.com/crawling/bots.json.

Are anthropic-ai and Claude-Web the same as ClaudeBot?

Prefer's free AI Crawler Access Checker covers the three bots Anthropic documents today: ClaudeBot, Claude-SearchBot and Claude-User. The names anthropic-ai and Claude-Web appear in some third-party lists but are not among the robots Anthropic's current documentation describes.

Questions

Asked plainly.

What does ClaudeBot do?

ClaudeBot is the crawler Anthropic uses to collect web content that could potentially contribute to training its generative AI models, and Prefer's free AI Crawler Access Checker shows whether your robots.txt lets it in. Anthropic says ClaudeBot helps enhance the utility and safety of those models, and that its bots honour industry standard robots.txt directives.

How do I block ClaudeBot?

Check what your robots.txt does today with Prefer's free AI Crawler Access Checker, or build the file with Prefer's free robots.txt generator. By hand, add 'User-agent: ClaudeBot' followed by 'Disallow: /' to the robots.txt in your top-level directory, and do it on every subdomain you want to opt out. Anthropic warns that blocking its IP addresses instead may not work reliably, because it stops ClaudeBot reading your robots.txt.

What is the difference between ClaudeBot and Claude-SearchBot?

ClaudeBot collects content that could contribute to model training, while Claude-SearchBot indexes content to improve search results for Claude users; Prefer's free AI Crawler Access Checker shows how your robots.txt treats each. Anthropic says disabling Claude-SearchBot may reduce your visibility in user search results. It documents no such effect for ClaudeBot.

Does blocking ClaudeBot stop Claude from citing my site?

Anthropic does not say so, and Prefer tracks Claude answers on its Enterprise plan if you need to watch them. Anthropic documents only that restricting ClaudeBot signals your site's future materials should be excluded from its training datasets. Search and user-requested visits run through Claude-SearchBot and Claude-User, which have their own tokens.

Sources

Where these facts come from.

Every fact on this page traces to one of these. Where Anthropic does not document something, the page says so.

  1. 01 Anthropic Help Center, web crawling and how to block the crawler Anthropic Help Center, web crawling and how to block the crawler The three Anthropic bots, their tokens and purposes, robots.txt and Crawl-delay support, the per-subdomain opt-out and the IP list. Vendor docs 1 Oct 2026
  2. 02 Anthropic, bot IP list Anthropic, bot IP list One combined IP list for all Anthropic bots, used for the verification step above. Vendor data 1 Oct 2026
  3. 03 Robots Exclusion Protocol, RFC 9309 The standard that defines how user-agent tokens, Allow and Disallow rules are interpreted. Crawl-delay is not part of it. Standard 1 Oct 2026

Get your free AI visibility report
in about 10 minutes.

See how answer engines describe your brand today, and where the openings are to outpace the competition.