Glossary
CCBot Common Crawl's open-dataset crawler.
The bot behind the open web archive that many AI models learned from. Prefer's free crawler checker shows whether your robots.txt lets it in, and this page covers what blocking it does and does not change.
CCBot is Common Crawl's crawler for an open web dataset that is widely used to train AI models; you control it in robots.txt with the token CCBot. Prefer's free checker tests your rules for it.
Key facts
At a glance.
The entity facts an assistant lifts first, each from Common Crawl's own pages.
- Operator
- Common Crawl Foundation · commoncrawl.org
- Type
- Web crawler for an open dataset, widely used as AI training data
- User agent token
- CCBot
- Full UA string
- CCBot/2.0 (https://commoncrawl.org/faq/)
- Respects robots.txt
- Yes, including Crawl-delay and nofollow
- IP ranges
- Published at index.commoncrawl.org/ccbot.json
- Reverse DNS
- *.crawl.commoncrawl.org on IPv4 (not yet over IPv6)
- Runs JavaScript
- No. JavaScript is not executed and cookies are not used
- First announced
- Not documented (Common Crawl was founded in 2007; data collected since 2008)
- Affects AI citations
- Not documented. Common Crawl has no answer product
How it works
What happens when CCBot visits.
Three steps. The third is why CCBot matters for AI: Common Crawl publishes the data, and model builders pick it up.
- 01
Reads robots.txt
Common Crawl says CCBot checks robots.txt first and only fetches a page if crawling it is allowed. It obeys Crawl-delay and honours nofollow, and it reads sitemaps listed in robots.txt.
- 02
Fetches the raw HTML
Allowed pages are fetched with plain HTTP GET requests. JavaScript is not executed and cookies are not used, so only the server-rendered HTML is collected.
- 03
Joins the open dataset
The crawl goes into Common Crawl's open repository, which anyone can access and analyse. Common Crawl says it has become one of the most widely used sources of training data for large language models.
Why it matters for AEO
Allow or block: the decision in one sentence.
Should you block CCBot?
For most brands that want AI visibility, no, and Prefer's free AI Crawler Access Checker shows what your robots.txt does with CCBot today. Common Crawl's dataset is widely used to train language models, so blocking it narrows how many future models learn your brand, while live AI citations come from each engine's own search crawler. Block it when the content is what you sell.
- You want future models to know your brand, products and category language
- Your pages are marketing, docs or editorial you already give away
- You are fine with an open archive of your public pages existing
- The content is the product: paid research, licensed data, a members archive
- Legal or licensing terms rule out inclusion in open datasets
- You block by path, so public pages still reach the dataset
Blocking stops future collection only. Common Crawl also keeps an Opt-Out Ledger for legal requests, but it asks publishers to use robots.txt first.
Allow or block it
Three robots.txt patterns that cover most cases.
Each pattern names the CCBot token explicitly. A wildcard Disallow catches CCBot too, which is how sites drop out of the dataset by accident.
robots.txtCopy the pattern that matches your decision above.
# Common Crawl's documented opt-out
User-agent: CCBot
Disallow: /
Common Crawl's own opt-out. Your live AI search crawlers are unaffected.
- Check for a wildcard firstA User-agent: * group with Disallow: / already blocks CCBot. A group that names CCBot overrides the wildcard for CCBot only.
- Blocking is not removalPages collected before the rule may already sit in published crawls. A Disallow prevents future collection only.
- Server-rendered HTML onlyCCBot does not run JavaScript, so content that only appears client-side is never collected, allowed or not.
User-agent: CCBot
Allow: /
Crawl-delay: 10
Stay in the dataset but slow the crawler down. Common Crawl says it obeys Crawl-delay.
- Check for a wildcard firstA User-agent: * group with Disallow: / already blocks CCBot. A group that names CCBot overrides the wildcard for CCBot only.
- Blocking is not removalPages collected before the rule may already sit in published crawls. A Disallow prevents future collection only.
- Server-rendered HTML onlyCCBot does not run JavaScript, so content that only appears client-side is never collected, allowed or not.
User-agent: CCBot
Allow: /
Disallow: /research/
Disallow: /members/
Keep public pages in the dataset, protect the paid archive.
- Check for a wildcard firstA User-agent: * group with Disallow: / already blocks CCBot. A group that names CCBot overrides the wildcard for CCBot only.
- Blocking is not removalPages collected before the rule may already sit in published crawls. A Disallow prevents future collection only.
- Server-rendered HTML onlyCCBot does not run JavaScript, so content that only appears client-side is never collected, allowed or not.
Verify a visit
How to confirm it was really CCBot.
Common Crawl warns that other crawlers falsely identify themselves as CCBot. It documents two checks beyond the user agent.
- 01
Match the user agent
Look for the CCBot token in the user agent field. The current string is CCBot/2.0 (https://commoncrawl.org/faq/); Common Crawl says it may increment the version number in the future.
- 02
Match the IP range
Compare the request IP against the ranges Common Crawl publishes at index.commoncrawl.org/ccbot.json. A CCBot user agent from an IP outside those ranges is a spoof.
- 03
Check reverse DNS on IPv4
A genuine IPv4 visit resolves to a hostname under crawl.commoncrawl.org. Common Crawl notes reverse DNS is not yet supported over IPv6, so use the IP list there.
203.0.113.42 - - [01/Oct/2026:09:14:07 +0000] "GET /pricing/ HTTP/1.1" 200 18422 "-" "CCBot/2.0 (https://commoncrawl.org/faq/)" In context
The term in a sentence.
01"We block CCBot on the members archive only, so our public guides still reach the open dataset that future models train on."
02"The log showed a CCBot user agent from an IP outside Common Crawl's published ranges, so we treated it as a spoof."
Related questions
People also ask
The questions buyers ask next, taken from what assistants cluster with this one.
Is CCBot an AI crawler?
CCBot builds an open web dataset rather than an AI product, but Common Crawl says that dataset has become one of the most widely used sources of training data for large language models. Prefer's free AI Crawler Access Checker groups it with the open training datasets for that reason. It does not fetch pages for live AI answers.
What is the CCBot user agent string?
At the checked-on date Common Crawl's FAQ gives it as CCBot/2.0 (https://commoncrawl.org/faq/), and Prefer's free robots.txt generator writes rules for its robots.txt token, CCBot. The older bot identified itself as CCBot/1.0, and Common Crawl says the version number may increase.
See how to verify a visit →Does blocking CCBot remove my pages from AI models?
Not from models that already trained on earlier crawls. Prefer's free AI Crawler Access Checker confirms the block is in place, but a Disallow only stops future collection, so pages already in published Common Crawl datasets stay there.
Questions
Asked plainly.
What does CCBot do?
CCBot is the crawler Common Crawl uses to build an open repository of web crawl data that anyone can access and analyse. Prefer's free AI Crawler Access Checker tests whether your robots.txt lets CCBot in today. Common Crawl does not train models itself, but it says its data has become one of the most widely used sources of training data for large language models.
How do I block or allow CCBot?
Prefer's free robots.txt generator includes CCBot by name, and its 'Block bulk datasets' preset blocks it while keeping the AI search crawlers allowed. By hand, Common Crawl's own instruction is to add 'User-agent: CCBot' followed by 'Disallow: /'. It also obeys the Crawl-delay parameter if you only want to slow it down.
Does blocking CCBot stop ChatGPT or Perplexity from citing my site?
Common Crawl does not document any effect on live AI answers, and it has no answer product of its own. Prefer tracks citations on ChatGPT, Gemini, Perplexity, Google AI Overviews and AI Mode, and those engines fetch live pages through their own crawlers, not CCBot. Blocking CCBot is a decision about future training datasets, not about this week's citations.
Does CCBot run JavaScript?
No, and Prefer's free AI Crawler Access Checker only covers the robots.txt side of that question. Common Crawl's FAQ says JavaScript is not executed and cookies are not used, so CCBot only sees the HTML your server returns, and content that only appears after JavaScript runs is never collected.
Sources
Where these facts come from.
Every fact on this page traces to one of these. Where Common Crawl does not document something, the page says so rather than guessing.
- 01 Common Crawl, CCBot What CCBot is for, in Common Crawl's own words. Vendor docs 1 Oct 2026
- 02 Common Crawl, FAQ The robots.txt token, the user agent string, Crawl-delay support, the opt-out lines and the no-JavaScript note. Vendor docs 1 Oct 2026
- 03 Common Crawl, CCBot IP ranges The published IP list used for the verification step above. Vendor data 1 Oct 2026
- 04 Common Crawl, About The founding date and the statement that the dataset is widely used to train large language models. Vendor docs 1 Oct 2026
- 05 Common Crawl, Opt-Out Registry announcement The Opt-Out Ledger for legal requests, and the advice to use robots.txt first. Vendor blog 1 Oct 2026
CCBot and Common Crawl are names of the Common Crawl Foundation. Prefer is not affiliated with, endorsed by or sponsored by Common Crawl. This entry reflects public documentation at the dates shown. Something wrong here? Tell us and we will fix it →
Keep reading