Glossary
Crawl Budget how much of your site a crawler fetches.
Crawl budget is the set of URLs a search engine can and wants to crawl on your site. When it matters, how AI crawlers fit in, and how Prefer helps.
Crawl Budget is the set of URLs a search engine can and wants to crawl on your site, set by your server's capacity and its demand. Prefer's Agent Analytics shows which AI crawlers visit.
Crawl budget is the set of URLs a search engine can and wants to crawl on your site, limited by how much load your server can take and how much the engine wants your pages. Prefer’s Agent Analytics reads your server logs to show which AI crawlers visit and what they fetch, the AI side of the same question. For most small and mid-size sites, crawl budget is not the bottleneck; access is.
How crawl budget works#
Google’s guide to managing crawl budget defines it as “the set of URLs that Google can and wants to crawl,” made of two parts:
- Crawl capacity limit. How much Googlebot can fetch without overloading your server. Google says that if your site responds quickly and steadily the limit goes up, and if it slows down or returns server errors (5xx) or rate-limiting signals (HTTP 429), the limit goes down.
- Crawl demand. How much Google wants your pages, based on things like site size, how often pages change, page quality and relevance.
Google is explicit about who needs to care. The guide is written for large sites with over a million unique pages that change about weekly, and for sites with over 10,000 pages that change daily.
Why it matters for AI search visibility#
A page that is never fetched cannot be indexed, and a page that is not indexed or retrieved cannot be cited in an AI answer. On a large site, crawlers spending their time on duplicate filter pages or dead URLs means your important pages get fetched less often, so the version an engine has may be stale.
Google’s guide covers Google’s crawlers. AI companies run their own bots, each documented separately. OpenAI, for example, lists its bots: GPTBot collects training data, OAI-SearchBot surfaces sites in ChatGPT search, and ChatGPT-User visits pages on a user’s behalf. Do not assume Google’s crawl budget numbers apply to them. Your server logs show what each AI crawler actually requests.
Common confusions#
Crawl budget is not a ranking factor. Being crawled more often does not make a page rank or get cited more. It only means the engine has a fresher copy.
Blocking is not budgeting. Disallowing a bot in robots.txt removes those pages from what that bot fetches. Block OAI-SearchBot and, per OpenAI, your site will not appear in ChatGPT search answers. Use robots.txt to cut waste, not to cut off the crawlers you want citing you. The should I block AI crawlers answer covers the trade-offs.
Example#
An online store has 5,000 products, but filter combinations (colour, size, price) create 800,000 crawlable URLs. Crawlers spend most of their visits on near-duplicate filter pages, and new product pages wait days to be fetched. Blocking the filter parameters in robots.txt and keeping the sitemap current points crawlers back at the pages that matter.
How to protect your crawl budget#
Fig. 01
Four ways to stop wasting crawls
How to check crawl activity#
For Google, the Crawl Stats report in Search Console shows how many requests Googlebot made, how fast your server responded and which status codes it returned. For AI crawlers, read your server logs and filter by user-agent, such as GPTBot, OAI-SearchBot or ClaudeBot. Look for three things: whether the bot visits at all, which pages it fetches, and how many of its requests hit errors or redirects.
How Prefer helps#
Prefer’s Agent Analytics reads your server logs to show which AI crawlers visit and what they fetch, and the free Crawler Access Checker tests whether your robots.txt lets the AI search crawlers in. Then Prefer tracks whether five AI engines cite you.
In context
The term in a sentence.
Related questions
People also ask.
- Does crawl budget matter for small sites?
- How do I check my crawl budget?
- Do AI crawlers have a crawl budget?→
Questions
Asked plainly.
What is crawl budget in simple terms?
It is how many pages a search engine can and wants to fetch from your site in a given period. Prefer's Agent Analytics reads your server logs to show which AI crawlers visit and which pages they fetch. Google defines it as the set of URLs Googlebot can and wants to crawl.
Does crawl budget matter for my site?
For most sites, no, and Prefer's free Crawler Access Checker covers the check that matters more for them: whether AI crawlers are allowed in at all. Google's own guide is written for sites with over a million unique pages, or over 10,000 pages that change daily. Below that, Google's crawlers can usually fetch everything you publish.
Do AI crawlers like GPTBot have a crawl budget?
Every crawler limits how much it fetches, but Google's crawl budget guide describes Google's crawlers, not OpenAI's, Anthropic's or Perplexity's. Prefer's Agent Analytics shows what AI crawlers actually request from your server, which is the reliable way to see their behaviour on your site. OpenAI notes it can take about 24 hours for its systems to adjust after a robots.txt change.
How do I stop wasting crawl budget?
Follow Google's advice: consolidate duplicate content, block unimportant URLs in robots.txt, return 404 or 410 for removed pages, fix soft 404s and keep sitemaps current. Prefer's free robots.txt generator helps write the rules, and its Crawler Access Checker confirms you have not blocked the AI search crawlers by mistake.
Keep reading