Glossary

AI Crawler the bots that fetch your pages for AI.

  • Technical

An AI crawler is a bot that fetches your pages to train models or source AI answers. The main ones, their user-agents, and how to check yours with Prefer.

Updated19 Sept 2026
Definition37 words

AI Crawler An AI crawler is a bot AI companies use to fetch web pages, to train models or to source live answers, named by a user-agent such as GPTBot. Prefer's free checker shows which ones can reach you.

Related terms ↓

An AI crawler is a bot that AI companies use to fetch web pages, either to train their models or to gather live sources for an answer, each identified by a user-agent such as GPTBot or ClaudeBot. Prefer’s free AI Crawler Access Checker shows which of them can reach your site. A crawler is the on-ramp to AI visibility: if a crawler cannot read your page, that engine cannot learn from it or cite it. Crawlers are controlled the same way search crawlers always have been, through your robots.txt file.

How an AI crawler works#

An AI crawler requests your pages over HTTP, identifies itself with a user-agent string, and, if your robots.txt allows it, reads what your server returns. One detail decides most outcomes: AI crawlers generally do not run JavaScript. They read the raw HTML your server sends, so any content that only appears after the page loads in a browser is invisible to them. That is why server-rendering matters more for AI visibility than almost any other technical fix.

A robots.txt allow-list permitting 5 AI crawlers, each with an Allow check: GPTBot (OpenAI (training)), OAI-SearchBot (ChatGPT Search), ClaudeBot (Anthropic (training)), PerplexityBot (Perplexity), Google-Extended (Gemini (training and grounding)). User-agents current as of September 2026; confirm each against the vendor's own crawler documentation. Allowing a crawler only makes you eligible. It does not guarantee a citation.
A robots.txt allow-list for the major AI crawlers, each with the engine it serves. Allowing access is necessary but not sufficient for a citation. Check your crawler access

Two jobs: training versus live fetch#

Not all AI crawlers do the same thing, and the difference changes what blocking one costs you. Training crawlers build a model’s baked-in knowledge; live-fetch crawlers pull sources at the moment an assistant answers. Blocking a training crawler removes you from future models’ memory. Blocking a live-fetch crawler makes you ineligible to be cited in that engine’s real-time answers.

Two columns. Training crawlers, build the model's memory: GPTBot, ClaudeBot, Collect pages into a training corpus, Blocking removes you from future models, No immediate traffic either way. Live-fetch crawlers, source the answer in real time: OAI-SearchBot, PerplexityBot, Fetch pages when an assistant answers, Power the inline citations you can win, Can send AI referral traffic.
AI crawlers split into two jobs: training the model versus fetching live sources for an answer. Blocking each one costs you something different. Conceptual comparison.

Why AI crawlers matter for your brand#

A crawler is the difference between being readable and being invisible to an entire engine. ChatGPT alone reaches your content through more than one route: a frozen training corpus (via GPTBot), live browsing on a search index, and its own dedicated search crawler (OAI-SearchBot) that powers inline citations. Understanding which route matters helps you spend effort where it moves the number.

A flow diagram. Your site, drawn as a small browser window with an answer-first passage highlighted, feeds three crawler routes: the frozen training corpus (GPTBot), from which ChatGPT answered most buyer questions in our July 2026 study, citing nothing; live browsing, widely reported to use Bing's index (Bingbot), which OpenAI does not document; and ChatGPT Search (OAI-SearchBot), highlighted as your biggest lever, a dedicated index that crawls the web continuously and powers the inline citations. The routes converge into a ChatGPT answer that names your brand with citation 1 pointing at yoursite.com.
One brand can reach a ChatGPT answer through three crawler routes: the training corpus (GPTBot), live browsing on a search index, and ChatGPT Search (OAI-SearchBot), which powers the inline citations. Conceptual illustration.

What an AI crawler does, at a glance#

Summary graphic of 4 items: 1. Fetches raw HTML: Reads what your server returns; it does not run client-side JavaScript. 2. Announces itself: Identifies itself with a user-agent so you can allow or block it by name. 3. Respects robots.txt: The major crawlers obey robots.txt rules; that file is your control panel. 4. Eligibility, not a promise: Access lets you be read and cited; it never guarantees the citation.
Four things to know about AI crawlers: they read raw HTML, name themselves, obey robots.txt, and grant eligibility rather than guaranteed citations. Conceptual summary.

Example#

Say you allow every AI crawler but your product pages render their content with client-side JavaScript. GPTBot and OAI-SearchBot request the page, receive a near-empty HTML shell, and move on. You did everything right in robots.txt and still cannot be cited, because the crawler never saw your words. Switch those pages to server-rendered HTML and the same crawlers now read the full content, making you eligible to appear in answers.

How Prefer helps with AI crawlers#

Prefer’s audit checks whether AI crawlers can actually reach your site, whether your key pages are server-rendered so those crawlers can read them, and, far more importantly, whether the engines then cite you, because crawlability is the floor, not the finish line. Our free AI Crawler Access Checker tests your robots.txt in seconds. Run a full AI visibility audit for the whole picture.

In context

The term in a sentence.

Related questions

People also ask.

Questions

Asked plainly.

What is an AI crawler in simple terms?

A bot that AI companies send to read your website. Prefer's free AI Crawler Access Checker shows which AI crawlers your robots.txt lets in. Some crawlers collect pages to help train a model; others fetch pages live when an assistant is answering a question. Each announces itself with a user-agent name, like GPTBot for OpenAI or ClaudeBot for Anthropic, so you can allow or block it in robots.txt.

What are the main AI crawlers?

The most common are GPTBot and OAI-SearchBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot (Perplexity), and Google-Extended (Google). Prefer's free AI Crawler Access Checker shows which of these your robots.txt lets in. GPTBot and ClaudeBot gather training data; OAI-SearchBot and PerplexityBot fetch live sources that power inline citations. Google-Extended is not a separate crawler but a robots.txt token: it controls whether Google may use your pages to train Gemini models and to ground answers in Gemini Apps and Vertex AI, and it does not affect Google Search.

Should I block or allow AI crawlers?

For most brands that want AI visibility, allow them. Prefer's free AI Crawler Access Checker shows which ones your robots.txt blocks today. Blocking a training crawler removes you from future models' knowledge, and blocking a live-fetch crawler makes you ineligible to be cited in that engine's answers. Blocking is a legitimate choice for privacy or licensing reasons, but understand it is a visibility trade-off, not a neutral default.

Do AI crawlers run JavaScript?

Generally no, and Prefer's audit checks whether your key pages are server-rendered so these crawlers can read them. Most AI crawlers read the raw HTML your server returns and do not execute client-side JavaScript, so content that only renders in a browser is invisible to them. Server-rendering your pages is the single most important technical step for AI crawlability.

Get your free AI visibility report
in about 10 minutes.

See how answer engines describe your brand today, and where the openings are to outpace the competition.