Earn the editorial
40% of citations sit with independent writers. Getting into the roundups and guides that already rank for your category is the single highest-leverage move available.
The Prefer Index · Edition 03 · Q3 2026
Three times a year we run the same 1,237 buyer questions across every major answer engine and record every source behind every answer. This is the third edition: 18,412 citations, 713 domains, and a clear picture of who AI trusts.
Key findings
Independent editorial sites take 40% of all citations in AI answers, more than four times what vendor-owned pages receive. Across 1,237 buyer questions and five engines, ten domains carried 31% of every citation, and no two engines ever produced the same source set for the same question. For brands, the practical finding is that most AI visibility is earned on other people's websites.
Vendor-owned pages take 9%. The single largest source of AI recommendations is writing you do not control and cannot buy.
713 domains were cited at least once. The top ten accounted for 31% of all 18,412 citations, and the top fifty for 58%.
Across 6,185 answers, there was not one question where all five engines produced the same set of sources. Only three domains were cited by all five engines anywhere.
For X vs Y questions, 82% of answers cited neither vendor. The engines went looking for someone impartial and found one.
44% of brands named in an answer were never the source behind it. Mentions and citations move independently, and only one of them is directly earnable.
The dataset
Across 8 categories, from category discovery to pricing and integration
ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews
Every question on every engine, captured in full with its sources
Each traced to a domain, a URL and a source type
Every site any engine used at least once
Every product or company mentioned in any answer
Run 14 to 21 July 2026, from a clean US English context with no personalisation, no logged-in accounts and no prior conversation history.
01 · Where citations go
Every citation was classified by what kind of site it came from. The pattern is consistent enough across engines that it functions as a rule: AI builds answers from independent writing first, and vendor pages last.
Share of all citations
Citation share by source type
18,412 citations across five engines, classified by site type.
Per engine
The same split, engine by engine
Perplexity leans hardest on community. Gemini is the most willing to cite a vendor. Claude is the most editorial.
02 · How old the sources are
We resolved a publish or last-updated date for 14,806 of the 18,412 citations. AI is not reading this week's news, it is reading the archive, and the archive keeps working for years.
Age of cited page at time of citation
When the cited pages were published or last updated
14,806 citations with a resolvable date, bucketed by age.
The practical reading: a page published today is an asset that compounds. Most of what gets cited in 2028 is being written now.
03 · What format gets cited
Source type tells you whose site. Format tells you what to write. We classified every cited URL by what kind of page it actually is.
Share of citations by page format
Citation share by page format
All 18,412 citations, classified by page type rather than site type.
04 · The concentration curve
Citation is far more concentrated than organic search. A long tail exists, but it is thin, and the head is where the recommendations actually come from.
Cumulative citation share
How citation share accumulates by domain rank
Domains ranked by total citations, cumulative share of all 18,412.
05 · Engine personalities
Same questions, same week, five very different reading styles. How many sources each engine pulls, how widely it ranges, and how willing it is to cite a vendor.
Cites moderately and repeats its favourites. The narrowest domain range of the five, so being in its set matters more.
Cites twice as much as anyone else and ranges widest. The easiest engine to enter and the hardest to dominate.
Most willing to cite a vendor directly, and leans on Google's own surfaces more than the others.
Cites least and most selectively, with the strongest editorial and reference bias. Hardest set to enter.
Shortest answers, most brands named per answer. Optimises for a quick shortlist rather than an argument.
05b · The agreement problem
We checked how many domains each engine shared with the others. The overlap is far smaller than most teams assume, which is why a single-engine view of your visibility is close to worthless.
Domain overlap across engines
How many domains were cited by how many engines
Of 713 domains cited at least once, by number of engines citing them.
06 · How stable answers are
We re-ran a 200-question subset three times over four days on each engine and measured how much the sources and the named brands changed. This is the number that should govern how much you trust any single AI visibility reading, including ours.
Brands are far stickier than sources. An engine will reach for a different article and still recommend the same three products, which is exactly why mention share and citation share have to be measured separately.
What this means for measurement: a single reading of your AI visibility carries meaningful noise. A trend across weekly readings does not. Any tool reporting a precise figure from one run is overstating its own precision.
07 · By question type
The 1,237 questions split into six families. Which sources an engine reaches for depends more on what kind of question it is than on which engine is answering.
Source mix by question family
Where citations come from, per question type
Six families across 1,237 questions. Percentages are of citations within that family.
08 · Most-cited domains
Every domain cited more than forty times, with which engines used it. Sorted by total citations across the full run.
Top 20 of 713 cited domains. The full ranked list ships with the dataset.
09 · By category
The overall split hides real variation. In regulated categories AI leans on institutions; in consumer categories it leans on communities. Where you need to show up depends entirely on which of these you sell into.
Source mix by category
Where citations come from, per category
Eight categories, each with its own retrieval habits. Percentages are of citations within that category.
10 · The Index
Share of voice across every answer in the run: how often each brand was named as a proportion of all 2,904 brand mentions. This is the Index itself, and it is a measure of recommendation, not of market share.
Share of all brand mentions
Most-recommended brands across all categories
Of 2,904 brands named across 6,185 answers, top 14 by mention share.
11 · How brands are described
Every brand mention was classified by how the surrounding sentence described it. Most mentions are neutral, but the critical ones cluster around a small number of recurring themes.
Tone of brand mentions
How AI describes the brands it names
All 2,904 brand mentions across 6,185 answers, classified by sentence-level tone.
Caveat themes
What the caveats are about
Of every qualified or critical mention, what the reservation was about.
Nearly every critical mention traced back to a review platform or forum thread rather than to editorial writing. Sentiment is repaired where the reviews live, not on your own site.
12 · Comparison questions
We isolated the 214 head-to-head comparison questions in the set. These are the highest-intent questions a buyer asks, and they are the ones brands lose most decisively.
The engine went looking for an impartial account and found one somewhere else.
Almost always the vendor with a genuinely balanced comparison page, not a sales page.
G2, Capterra and TrustRadius are the default fallback when no neutral article exists.
Reddit outranked both vendors on 46 of the 214 questions.
The vendors that did get cited had one thing in common: a comparison page that named a competitor's advantage in plain language. Pages that only listed their own wins were not cited once.
13 · What changed
Compared against Edition 02, published March 2026, and Edition 01, published November 2025.
Across three editions
What moved between editions
Edition 01 November 2025, Edition 02 March 2026, Edition 03 August 2026.
Methodology
1,237 questions drawn from real buyer language across eight categories, weighted to match the distribution of commercial intent we see in customer prompt sets. The full list ships with the dataset.
Every question was submitted to all five engines between 14 and 21 July 2026, from a clean US English context: no account, no personalisation, no conversation history, fresh session per question.
Answers were recorded in full, along with every cited URL and every visible source panel. Where an engine cited the same domain twice in one answer, it was counted once.
Each domain was classified into one of six source types by two reviewers independently, with disagreements resolved by a third. Inter-rater agreement was 94%.
Product and company names were extracted from answer text and normalised against a 12,000-entry entity list. Ambiguous names were reviewed by hand.
No paid placements, no logged-in accounts, no regional variation, and no repeated sampling of the same question within the window. Absolute figures will differ from a run in another week; the shape should not.
Cite this research
The Prefer Index is published under CC BY 4.0. You are free to reproduce any figure or chart, commercially or otherwise, with attribution.
The Prefer Index, Edition 03. Prefer, August 2026. https://tryprefer.com/data/benchmark/
The questions readers ask next, taken from what assistants cluster with this one.
Reddit, G2 and Wikipedia lead, and they are the only three domains cited by all five engines tested. Reddit alone accounted for 5.7% of all 18,412 citations recorded, more than any vendor-owned domain in the entire dataset.
Questions
A recurring benchmark of how AI answer engines cite sources and recommend brands. Three times a year we run the same 1,237 buyer questions across ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews, record every answer and every citation, and publish the full dataset.
Because AI answers are assembled from a handful of retrieved sources rather than ranked from an index. Which kind of site an engine reaches for determines who gets named, so understanding the mix tells you where to invest before it tells you how to optimise.
Yes. The full dataset, all 18,412 citation rows with question, engine, answer text and classified source, is available as a CSV under a permissive licence. Attribution is appreciated, and a link back helps the next edition.
Please do. Cite it as The Prefer Index, Edition 03, Prefer, August 2026, and link to this page. If you are writing about a specific figure, the dataset lets you verify it yourself.
Question set, engine mix and classification method all vary between studies, and answer engines are non-deterministic. Absolute percentages should be treated as approximate. Direction and relative ordering are the durable findings.
A ranking study measures positions in a list of links. This measures which sources an engine chose to build an answer from, and which brands it named as a result. The two correlate loosely and diverge often.
Edition 04 will add Grok and Copilot, and we are evaluating regional engines. The core five will stay constant so the time series remains comparable.
Yes, that is what the product does. A Prefer workspace runs your own prompt set continuously rather than three times a year, sliced by your competitors, markets and buyer personas.
What it means
Nine in ten citations point somewhere you do not control. That is not a reason to stop publishing, it is a reason to be honest about what publishing does: it makes you quotable once someone else has already put you in the conversation.
40% of citations sit with independent writers. Getting into the roundups and guides that already rank for your category is the single highest-leverage move available.
16% of citations come from review platforms, and they are the fallback whenever no neutral article exists. Placement, category tags and review volume all read straight into answers.
The only vendor pages cited on comparison questions were the ones that named a competitor's advantage. Every purely promotional comparison page in the set was ignored.
Community share is rising faster than any other source type. The archives being retrieved today were written years ago, which is exactly why starting now matters.
See how answer engines describe your brand today, and where the openings are to outpace the competition.