Every answer lists crawl access; none names ChatGPT’s search crawler#
Prefer asked ChatGPT, Claude, Gemini and Perplexity this question three times each on 13 September 2026.
| What the answer listed | Answers (of 12) |
|---|---|
| Crawl access (robots.txt, indexing) | 12 |
| A direct answer at the top | 12 |
| Schema markup, always including FAQPage | 12 |
| Named authors or E-E-A-T signals | 11 |
| Tracking AI answers and citations | 10 |
| Original or primary-source data | 9 |
| Consistent brand facts across the web | 8 |
| Mentions, reviews or PR on other sites | 8 |
| Presence on Reddit or forums | 3 |
| llms.txt | 1 |
The five answers that named crawlers all listed GPTBot and ClaudeBot, which OpenAI and Anthropic describe as training crawlers. None of the 12 named OAI-SearchBot or Claude-SearchBot, the crawlers OpenAI and Anthropic document for search.
Eight checks, in order#
Items 1 and 2 are crawler access, 3 to 5 the page, 6 measurement, 7 and 8 other sites. The first two can block a citation outright.
- Search crawlers can reach you. Done when the crawler access checker shows OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot allowed on public pages. Allow Google-Extended too if you want Gemini app answers to use your pages; it does not affect Google Search.
- The answer is in the HTML. ChatGPT’s and Claude’s crawlers did not run JavaScript in Vercel’s December 2024 analysis. Done when a key page shows the answer with JavaScript off.
- The answer comes first. Every answer lists it, so it is the minimum. Done when each priority page’s first two sentences answer its question and make sense alone.
- Every claim has a source. Original data made 9 lists, none of them Perplexity’s. Done when each number has a source and a date, the author is named, and something is first-hand: your data, a test or an example.
- Schema matches the page. The next section covers what it no longer earns. Done when the Schema Markup Validator (validator.schema.org) shows no errors and nothing is marked up that a reader cannot see.
- A fixed prompt set, run more than once. Run 10 to 50 buyer prompts (how many to track) on each engine. None of the 12 answers said to repeat a prompt within a check, though answers change between runs. Done when each prompt has three or more runs per engine and you re-run the set weekly (the weekly loop).
- Your facts match everywhere. Done when your one-line description matches on your site, LinkedIn, G2 and Crunchbase.
- You are in the sources engines cite. Only Gemini’s three answers listed presence on Reddit or forums, yet in our August study Reddit was the most-cited source in ChatGPT answers on all 25 software buying queries (one run on one day, so directional). Done when each source from item 6 that leaves you out has an owner and a next step (where to get listed).
Schema and llms.txt matter less than the lists suggest#
All 12 answers named FAQPage markup, and none mentioned that Google stopped showing FAQ rich results on 7 May 2026. Ten named HowTo, whose rich results Google dropped in September 2023. The markup can still label your content for engines; only Google’s rich result display is gone.
llms.txt made 1 of 12 lists. Google says AI Overviews need no AI text files, yet its AI Overview for this question told readers to add one on 13 September (is llms.txt worth it?).
The most-cited checklist names a crawler Anthropic no longer uses#
AirOps’s “AEO Checklist: 48 Critical Factors”, on LinkedIn or its blog, was a source in 6 of the 12 answers. It rightly puts crawl access first and lays out a 90-day plan in three stages. Its crawler item names “GPTBot, Claude-Web, or PerplexityBot”, but Anthropic told 404 Media in July 2024 that Claude-Web was no longer in use. One Claude answer repeated the AirOps sentence word for word. Google’s AI Overview, citing savvy.co.il and the LinkedIn post, listed the same three bots; AI Mode, also citing the post, named GPTBot, OAI-SearchBot, PerplexityBot and Google-Extended.
Disclosure: Prefer sells AI visibility tracking (item 6), and AirOps sells in the same market. For a baseline, run a free AI visibility audit.
Sources
- Across 12 answers to 'What should be on an AEO checklist?' from ChatGPT, Claude, Gemini and Perplexity (three runs each, 13 September 2026), counted from the raw answer text: crawl access (robots.txt or indexing) 12, a direct answer at the top 12, schema markup 12 (all 12 named FAQPage, 10 named HowTo), named authors or E-E-A-T signals 11, tracking AI answers and citations 10, original or primary-source data 9 (none from Perplexity), consistent brand facts across the web 8, mentions, reviews or PR on other sites 8, presence on Reddit or forums 3 (all Gemini), llms.txt 1 (Claude). None of the 12 said to repeat a prompt within a check, and none mentioned that Google no longer shows FAQ or How-to rich results. (Prefer research note: AEO checklist (2026-09-18), dated research note)
- Five of the 12 answers named specific crawlers to allow (two Claude, three Gemini). All five named GPTBot, ClaudeBot and PerplexityBot; three also named Google-Extended and one named Claude-Web. None of the 12 named OAI-SearchBot or Claude-SearchBot. (Prefer research note: AEO checklist (2026-09-18), dated research note)
- Five answers set a 40 to 60 word length for the opening answer (one ChatGPT answer, for 'What is' definitions only; all three Claude answers; one Gemini answer) and two Gemini answers said the first one or two sentences. The three Claude answers used wording found verbatim in pages they cited (a geeks360.net checklist and an AirOps guide), and none of the five gave a measured result. All three ChatGPT answers said AEO does not replace SEO. All three Gemini answers gave the Authorized Economic Operator (customs) checklist first. (Prefer research note: AEO checklist (2026-09-18), dated research note)
- Prefer's crawler checks treat OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot as the search crawlers, separate from the training crawlers and controls GPTBot, ClaudeBot and Google-Extended (operator docs re-checked 13 September 2026). (Prefer evidence note: AI visibility funnel, crawler correction (2026-09-13), dated research note)
- OpenAI says GPTBot crawls content that may be used in training its models, OAI-SearchBot is used to surface websites in ChatGPT's search features, sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers (though they can still appear as navigational links), and each setting is independent of the others. OpenAI: Overview of OpenAI crawlers
- Anthropic documents three agents: ClaudeBot, which collects content that could contribute to model training; Claude-User, for user-initiated requests; and Claude-SearchBot, which improves search result quality. The page (last updated April 2026, checked 18 September 2026 in a rendered browser) does not list Claude-Web. Anthropic Help Center: Does Anthropic crawl data from the web
- An Anthropic spokesperson told 404 Media in July 2024 that the ANTHROPIC-AI and CLAUDE-WEB user agents are no longer in use. An update dated 30 July 2024 adds that ClaudeBot will respect blocks aimed at those two older agents. 404 Media: Websites are blocking the wrong AI scrapers (29 July 2024)
- Google-Extended controls whether Google-crawled content is used to train Gemini models and to ground answers in Gemini Apps and Vertex AI, and does not affect a site's inclusion in Google Search. Google Search Central: Google's common crawlers
- Google says a page must be indexed and eligible for a snippet in Google Search to appear as a link in AI Overviews or AI Mode, with no additional requirements, no AI text files and no special schema.org structured data needed. Google Search Central: AI features and your website
- Google stopped showing FAQ rich results in Search from 7 May 2026, and stopped showing How-to rich results in September 2023. (Prefer verification note: Google FAQ rich result (2026-09-15), dated research note)
- Vercel's December 2024 analysis of AI crawler traffic found that ChatGPT's and Claude's crawlers do not execute JavaScript. Vercel: The rise of the AI crawler
- In Prefer's August 2026 study of 25 B2B software buying queries (DataForSEO LLM Mentions, 19 August 2026), Reddit was the most-cited source in ChatGPT answers on all 25 queries, 30.3% of all source citations. One measured run on one day, so the figures are directional. (Prefer Reddit citation study methodology (2026-08-19), dated research note)
- AirOps's 'AEO Checklist: 48 Critical Factors' (LinkedIn, 18 March 2026; blog version dated 20 January 2026) was a source in 6 of the 12 answers (all three Perplexity answers cited the LinkedIn copy; two Claude answers and one Gemini answer cited the blog copy). It puts technical access first ('Technical access issues block everything else, so fix those first') and gives a roadmap in three 30-day stages. Its crawler item reads 'GPTBot, Claude-Web, or PerplexityBot'; one Claude answer repeated that sentence word for word. (Prefer research note: AEO checklist (2026-09-18), dated research note)
- On 13 September 2026, Google's AI Overview for 'What should be on an AEO checklist?' told readers to allow GPTBot, Claude-Web and PerplexityBot (citing savvy.co.il and the AirOps LinkedIn post) and to add an llms.txt file. Google's AI Mode, citing the AirOps post, named GPTBot, OAI-SearchBot, PerplexityBot and Google-Extended. (Prefer research note: AEO checklist (2026-09-18), dated research note)
- Prefer sells AI visibility tracking, so we have a commercial stake in measurement being on the checklist. Prefer pricing (our product)
People also ask
- What is on an answer engine optimization checklist?
- What should an AEO audit check?
- What should a GEO checklist include?
- How do I check if my site is ready for AI search?
- Which AEO checklist items matter most?