Your brand is invisible in AI search for a specific reason, and it is probably not the reason you are fixing.
Most teams that worry about brand visibility in AI search jump straight to publishing more content or tweaking robots.txt. But AI visibility fails in stages, the way a sales funnel fails in stages, and the effort only pays off if it lands on the stage where you actually drop out. A brand that is crawlable, retrieved, and cited can still lose because the answer recommends a competitor. A brand that wins ChatGPT can still be invisible on the other three engines, and usually is.
This piece walks the funnel stage by stage with measured data at every step: two fresh studies we ran this week (a robots.txt audit of every domain ChatGPT cites in our category, and the same 15 prompts run through four engines on one day), our weekly tracking runs, our July four-engine citation study, and the Ahrefs AI Search Benchmark. The funnel framing itself has prior art: Maja Voje and Ahrefs published an AI visibility funnel infographic that got this right at the concept level. We have redefined the stages around what we can actually measure, and re-verified every number at its source.
One honest disclosure up front: we sell an AI visibility platform, so we have a stake in you caring about this. Every number below traces to a dated study you could rerun yourself, and none of them are about us.
What is the AI visibility funnel?#
The AI visibility funnel is a five-stage diagnostic that locates where your brand drops out of AI answers. The stages are sequential: fail one and the later stages never happen for that answer. Prefer’s tracking shows the later stages (cited, named, kept) on your own prompts across five AI engines.
- Reachable. AI crawlers can fetch your pages.
- Retrieved. The engine runs a live search on the question and pulls your page into the answer’s working set.
- Cited. Your page becomes one of the sources the answer is built on.
- Named. The answer actually recommends your brand.
- Kept. The win holds across engines, and week after week.
The funnel is a diagnostic, not a to-do list. You do not work all five stages. You find the one leaking worst and fix that. The rest of this article gives you the numbers to make that diagnosis.
Stage 1: Reachable. Can AI crawlers fetch your pages?#
Almost every page AI already cites is easy to crawl. This week we fetched and parsed robots.txt for all 136 domains our weekly ChatGPT tracking has ever seen cited in our category, and checked each against 14 named crawler tokens using the same rules the engines follow (longest-match wins, Allow beats Disallow on ties, no rule means allowed). The result: 134 of 136 domains (98.5%) allow all five search crawlers, meaning OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, and Bingbot, the bots each operator documents for its search answers. Only reddit.com and theatlantic.com block any of them.
The blocks that do exist target training, not answers. Bytespider (ByteDance) is blocked by 14 of the 136 domains, CCBot (Common Crawl) by 10, Google-Extended by 7, Applebot-Extended by 6, and GPTBot and ClaudeBot, the OpenAI and Anthropic training crawlers, by 5 each. Compare that with 2 blocks each for Claude-SearchBot and PerplexityBot, and 1 each for OAI-SearchBot, Googlebot, and Bingbot. Every training crawler is blocked more often than any search crawler, a pattern consistent with staying out of training datasets while staying in the answers.
The one exception: reddit.com blocks all 14 tokens in robots.txt and still got cited 18 times across three engines in our cross-engine run this week. Its content reaches the engines through licensing deals, not crawling. Unless you have a licensing deal with OpenAI, that is not your playbook. (More on Reddit’s outsized role in our Reddit citation study.)
What this stage means in practice: being crawlable is table stakes, and you should verify it rather than assume it. Run your site through the crawler access checker, make sure your pages render their content as server-side HTML because AI crawlers mostly do not run JavaScript, and stop there. Passing stage 1 earns you nothing. Failing it silently disqualifies you from everything that follows. If you are already cited and considering blocking training bots, the data above says the winners do exactly that without losing answer visibility.
Stage 2: Retrieved. Does the engine actually search for you?#
An AI engine can only cite your page in an answer it researched. When the model answers from memory (ungrounded, in the jargon), there are no sources at all, and stage 3 never happens. So the second gate is whether the engine runs a live web search for the questions you care about.
The good news: on real buyer questions, they almost always do now. In our four-engine run this week (15 prompts from our tracked portfolio, one day, web search enabled), ChatGPT, Perplexity, and Claude ran live searches on 15 of 15 prompts each, and Gemini on 13 of 15. That is 58 of 60 answers built on live retrieval. Our weekly ChatGPT tracking says the same: 47 of 50 prompts triggered live search on September 2, and 48 of 50 on September 7. The prompts that skip search are short definitional ones, where the model already knows the answer. (Full breakdown on whether ChatGPT uses live search.)
The caution: this gate moves. In our July study, ChatGPT on gpt-4o skipped live search on 29 of 47 buyer questions, searching on just 18. Two months and a model generation later, the same engine searches on nearly everything we ask it. The retrieval gate depends on the engine, the model version, and the shape of the question, which is why you measure it on your own prompts instead of trusting anyone’s average, ours included.
There is also a subtlety inside retrieval: engines rewrite your question into several searches of their own (called query fan-out), then read what those searches surface. You are not optimizing for the buyer’s phrasing, you are optimizing for the queries the engine derives from it. Pages that answer one narrow question directly, with the question phrasing in the heading, get pulled into more of those derived searches.
Stage 3: Cited. Whose pages actually power the answers?#
Once an engine searches, it picks a handful of pages to build from. This is the stage where most competitive visibility is won or lost, and the source mix is knowable.
We read and classified all 134 unique URLs ChatGPT cited across one 50-prompt week of our tracking. The mix: 67 of 134 (50.0%) were vendor-owned pages, 39 (29.1%) documentation and reference, 17 (12.7%) independent editorial, 7 (5.2%) review-site profiles, 3 (2.2%) community threads, and 1 (0.7%) seeded content. Half of what ChatGPT reads in this category is pages the vendors themselves wrote: comparison pages, pricing pages, docs. Owned pages are not brochures in AI search. They are retrievable answers.
The mix flips with the question. In our July study (47 buyer questions, four engines, 1,237 citations logged), independent roundups took 40% of citations and vendor sites 34%. And on head-to-head “X vs Y” questions, vendor pages fell to 18% of citations while neutral comparisons dominated. The narrower and more comparative the question, the more engines reach for a referee instead of a contestant. (Full study: who gets cited in AI search.)
Two findings from third parties sharpen this stage. The Ahrefs AI Search Benchmark (Q4 2025 to Q1 2026) found that 28% of ChatGPT’s 1,000 most-cited pages have zero organic visibility in Google, so AI citation is not a rankings echo; the engines read pages Google ignores. The same report found 53.4% of pages cited in AI Overviews are under 1,000 words, with essentially no correlation between length and citation. Short, direct, answer-shaped pages get cited. Word count is not a moat.
What to do at this stage is the heart of on-page AEO and off-page AEO: make your own comparison and pricing pages retrievable answers, and earn your way into the small set of independent reviews and roundups the engines keep reusing. In our tracking for the week of September 2, 2026, one tech-press site was cited in 12 of the 46 live-search answers that listed sources, across five of its pages; its review of one tool appeared in 7 of those answers. One well-placed site can reach a quarter of a category’s answers.
Stage 4: Named. Does the answer recommend you?#
Getting cited and getting recommended are different events. A mention is your brand named in the answer text. A citation is your page used as a source. They fail independently, and mixing them up is the most common measurement mistake in this discipline.
A brand with strong memory presence gets mentioned even when its pages are never retrieved, because the model absorbed years of coverage about it. A brand with well-shaped pages gets cited on facts while the recommendation goes to a rival with the stronger reputation. These need different fixes. Mentions respond to off-page work, being talked about in the places engines learn from. Citations respond to on-page work, having the retrievable page. Share of voice is the honest metric for the first; citation share per prompt for the second.
On the off-page side, the strongest external signal Ahrefs found for AI brand visibility is not backlinks. Across 75,000 brands, YouTube mentions correlated 0.737 with AI brand visibility, the strongest of every factor they tested, ahead of branded search volume and link metrics. That matches what we saw in our own August study, where YouTube was the single most-cited source on Google’s AI surfaces. The engines learn who exists partly from video, podcasts, and community talk, which is why authentic presence there moves stage 4 even when it never links to you.
Stage 5: Kept. Does the win survive engines and weeks?#
Here is the stage almost nobody measures, and it is the most brutal one. Visibility does not transfer between engines. This week we ran the same 15 prompts through ChatGPT, Perplexity, Gemini, and Claude on the same day. The four engines cited 255 unique domains between them. 179 of those domains (70.2%) were cited by exactly one engine. Only 4 of 255 (1.6%) were cited by all four. On 14 of the 15 prompts, the engines answering the same question shared zero source domains.
This is not a quirk of our sample. Our July study found the same shape at larger scale: 5 of 713 domains cited by all four engines, with roughly 75% single-engine. Ahrefs found that even inside Google, AI Mode and AI Overviews share only 13.7% of cited URLs, and that different engines gave word-for-word identical answers to the same query just 0.51% of the time. Four engines means four scoreboards.
The engines are also different sizes of gate. In our run, ChatGPT cited 2.9 sources per answer on average, versus 18.9 for Perplexity, 9.2 for Gemini, and 5.9 for Claude. Winning a Perplexity citation means being one of nineteen sources. Winning a ChatGPT citation means being one of three. Same question, completely different competitive game. (How each engine differs, and what to do about it, in win each engine.)
And “kept” has a second axis: time. The retrieval flip earlier in this article is the cautionary tale: the same engine went from searching on 18 of 47 questions in July to 47 of 50 in September, one model generation later. Answers built on live retrieval change as the retrieval changes, which is why a one-off audit goes stale fast. The only honest way to know your visibility is holding is to measure the same prompts on every engine on a schedule, which is what the weekly loop is for.
How do you improve brand visibility in AI search at each stage?#
Each stage has its own failure mode, so each has its own fix. The mapping:
- Losing at reachable? Fix robots.txt and server-rendered content once. Done when the crawler access checker shows all five search crawlers allowed on your money pages. Do not spend another hour here.
- Losing at retrieved? Reshape pages so each answers one question directly, in the phrasing buyers use. A page titled “Platform overview” becomes “What does [product] cost and what is included?”. Done when the engines’ answers to your tracked questions show live sources rather than memory.
- Losing at cited? Build the owned pages the data says get cited (comparisons, pricing, docs) and earn placements in the few roundups your engines keep reusing. That is on-page plus off-page AEO. Done when your URL shows up in the source list of answers to your money prompts.
- Losing at named? Work the surfaces engines learn brands from: video first (YouTube mentions were the strongest visibility correlate Ahrefs tested, at 0.737), current review-site profiles, and real answers in the communities your category argues in. Done when your share of voice on tracked prompts moves without your citation share moving, because that is the mention channel working.
- Losing at kept? Stop measuring one engine. Track your prompt portfolio across all four, weekly, and treat each engine as its own scoreboard. Done when the same prompt set has run on every engine for two straight weeks and you can see which wins held.
The funnel is the diagnosis. The AEO loop is the weekly operating process that runs that diagnosis continuously and turns it into a short action queue: measure the portfolio, classify the sources, read the winners, score honestly, act. The two are designed to snap together, and the instrumentation stays cheap: our 50-prompt ChatGPT week meters at $0.95, and the full four-engine run above cost $1.83.
If you want the diagnosis done for you, start with a free audit. It shows where AI engines name and cite you on real prompts, the same kind of measurement we run on ourselves.
How we measured this#
Two fresh studies, September 9, 2026. The robots study fetched robots.txt for all 136 domains our weekly ChatGPT tracking has seen cited and evaluated 14 crawler tokens per domain with standard robots semantics (named group resolution, longest-match, Allow wins ties; 130 of 136 served a readable file, the rest default to allowed). The cross-engine study ran a 15-prompt subset of our tracked portfolio through ChatGPT (gpt-5.6-luna), Perplexity (sonar-pro), Gemini (gemini-3.8-flash), and Claude (claude-sonnet-5) via the DataForSEO LLM Response API with web search enabled, 60 answers total, metered cost $1.83; Gemini’s redirect URLs were resolved to their real domains before any counting. Supporting data: our weekly tracking runs of 2026-09-02 and 2026-09-07, our July 10 four-engine study (47 questions, 1,237 citations), our August Reddit citation study, and the Ahrefs AI Search Benchmark Report Q4 2025 to Q1 2026 plus the Ahrefs AI traffic study, both re-verified at source. Single-day LLM runs are directional, not settled statistics: a rerun shifts specific domains and counts, and API endpoints proxy the consumer apps. We publish the pattern-level findings that survived two independent measurements months apart.
Correction, September 13, 2026. The first version of this article counted GPTBot and ClaudeBot as answer crawlers and reported 130 of 136 domains (95.6%) allowing all five. OpenAI and Anthropic document both as training crawlers; their search answers use OAI-SearchBot and Claude-SearchBot. Re-scoring the same September 9 data against each operator’s documented search crawler gives 134 of 136 (98.5%). The per-crawler block counts did not change, and the finding that cited sites block training bots rather than search bots got stronger.
People also ask
- What is the AI visibility funnel?
- How do I improve my brand's visibility in AI search?
- Why is my brand not showing up in AI search?
- Do AI engines use live web search?
- Does AI visibility on one engine carry to the others?