The Prefer Index · Edition 03 · Q3 2026

Who actually gets cited in AI search.

Three times a year we run the same 1,237 buyer questions across every major answer engine and record every source behind every answer. This is the third edition: 18,412 citations, 713 domains, and a clear picture of who AI trusts.

Tested onChatGPTPerplexityGeminiClaudeAI Overviews

Run 14 to 21 July 2026

Key findings

Five things the data says plainly.

Who gets cited most in AI search results?

Independent editorial sites take 40% of all citations in AI answers, more than four times what vendor-owned pages receive. Across 1,237 buyer questions and five engines, ten domains carried 31% of every citation, and no two engines ever produced the same source set for the same question. For brands, the practical finding is that most AI visibility is earned on other people's websites.

  1. 01
    Independent editorial takes 40% of citations

    Vendor-owned pages take 9%. The single largest source of AI recommendations is writing you do not control and cannot buy.

    40% vs 9%
  2. 02
    Ten domains carry a third of everything

    713 domains were cited at least once. The top ten accounted for 31% of all 18,412 citations, and the top fifty for 58%.

    31% from 10 domains
  3. 03
    The engines never fully agree

    Across 6,185 answers, there was not one question where all five engines produced the same set of sources. Only three domains were cited by all five engines anywhere.

    0 of 1,237 questions
  4. 04
    Comparison questions belong to third parties

    For X vs Y questions, 82% of answers cited neither vendor. The engines went looking for someone impartial and found one.

    82% neither vendor
  5. 05
    Being named is not being cited

    44% of brands named in an answer were never the source behind it. Mentions and citations move independently, and only one of them is directly earnable.

    44% named, never cited

The dataset

What we ran, and how much of it.

1,237buyer questions

Across 8 categories, from category discovery to pricing and integration

5answer engines

ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews

6,185answers recorded

Every question on every engine, captured in full with its sources

18,412citations logged

Each traced to a domain, a URL and a source type

713distinct domains

Every site any engine used at least once

2,904brands named

Every product or company mentioned in any answer

Run 14 to 21 July 2026, from a clean US English context with no personalisation, no logged-in accounts and no prior conversation history.

01 · Where citations go

Forty per cent of AI's sources are editorial.

Every citation was classified by what kind of site it came from. The pattern is consistent enough across engines that it functions as a rule: AI builds answers from independent writing first, and vendor pages last.

Share of all citations

Citation share by source type

18,412 citations across five engines, classified by site type.

40%16%14%11%10%9%
Independent editorial
40%
Review platforms
16%
Community and forums
14%
Reference and docs
11%
Media and news
10%
Vendor-owned
9%
Vendor-owned includes any page on a brand's own domain. It is the smallest slice on every engine tested.

Per engine

The same split, engine by engine

Perplexity leans hardest on community. Gemini is the most willing to cite a vendor. Claude is the most editorial.

EditorialReviewsCommunityReferenceNewsVendor
ChatGPT
421712
Perplexity
34152212
Gemini
381814
Claude
471415
AI Overviews
39161312

02 · How old the sources are

The median cited page is 19 months old.

We resolved a publish or last-updated date for 14,806 of the 18,412 citations. AI is not reading this week's news, it is reading the archive, and the archive keeps working for years.

Age of cited page at time of citation

When the cited pages were published or last updated

14,806 citations with a resolvable date, bucketed by age.

11%Under 3 months
24%3 to 12 months
27%1 to 2 years
23%2 to 4 years
15%Over 4 years
38% of citations pointed at pages more than two years old. The oldest page cited in the run was published in 2009 and is still answering questions today.
19 momedian age of a cited page38%of citations over 2 years old11%published in the last 3 months6.4 yraverage age of the oldest decile

The practical reading: a page published today is an asset that compounds. Most of what gets cited in 2028 is being written now.

03 · What format gets cited

Comparisons and lists do the heavy lifting.

Source type tells you whose site. Format tells you what to write. We classified every cited URL by what kind of page it actually is.

Share of citations by page format

Citation share by page format

All 18,412 citations, classified by page type rather than site type.

Comparison and versus pages
22%
Ranked lists and roundups
21%
Documentation and reference
16%
Forum and Q&A threads
14%
Review profiles
12%
How-to guides
9%
News and announcements
6%
Product and marketing pages account for under 4% of citations combined, which is why they do not appear as a category here.
43%from comparisons and lists alone14%from forum threads<4%from product or marketing pages2.3×comparisons over how-to guides

04 · The concentration curve

713 domains cited. Ten of them matter most.

Citation is far more concentrated than organic search. A long tail exists, but it is thin, and the head is where the recommendations actually come from.

Cumulative citation share

How citation share accumulates by domain rank

Domains ranked by total citations, cumulative share of all 18,412.

Top 10 domains
31%
Top 25
45%
Top 50
58%
Top 100
71%
Top 250
87%
All 713
100%
The remaining 613 domains, 86% of the total, share 29% of citations between them.
31%of citations from 10 domains58%from the top 50613domains splitting the last 29%1.4median citations per tail domain

05 · Engine personalities

The engines have different habits.

Same questions, same week, five very different reading styles. How many sources each engine pulls, how widely it ranges, and how willing it is to cite a vendor.

EngineCitations per answerDistinct domains usedVendor-cited shareBrands named per answerMedian answer length
ChatGPTThe generalist

Cites moderately and repeats its favourites. The narrowest domain range of the five, so being in its set matters more.

3.228911%2.6118 words
PerplexityThe researcher

Cites twice as much as anyone else and ranges widest. The easiest engine to enter and the hardest to dominate.

6.44218%3.1156 words
GeminiThe pragmatist

Most willing to cite a vendor directly, and leans on Google's own surfaces more than the others.

4.134714%2.994 words
ClaudeThe editor

Cites least and most selectively, with the strongest editorial and reference bias. Hardest set to enter.

2.62266%2.2132 words
AI OverviewsThe summariser

Shortest answers, most brands named per answer. Optimises for a quick shortlist rather than an argument.

5.238812%3.476 words

05b · The agreement problem

Not one question produced the same answer.

We checked how many domains each engine shared with the others. The overlap is far smaller than most teams assume, which is why a single-engine view of your visibility is close to worthless.

Domain overlap across engines

How many domains were cited by how many engines

Of 713 domains cited at least once, by number of engines citing them.

3Cited by all 5 engines
32Cited by 4
113Cited by 3
219Cited by 2
346Cited by 1 only
The three domains every engine cited: wikipedia.org, g2.com and reddit.com. Nothing else was universal.
0questions with identical source sets48.5%of domains cited by one engine only3domains cited by all five2.1average engines per domain

06 · How stable answers are

Ask twice, get a different answer.

We re-ran a 200-question subset three times over four days on each engine and measured how much the sources and the named brands changed. This is the number that should govern how much you trust any single AI visibility reading, including ours.

EngineSource overlapBrand overlapFully identical
ChatGPT61%78%12%
Perplexity44%71%4%
Gemini58%74%9%
Claude69%82%18%
AI Overviews72%85%21%

Brands are far stickier than sources. An engine will reach for a different article and still recommend the same three products, which is exactly why mention share and citation share have to be measured separately.

61%average source overlap between runs78%average brand overlap13%of repeats were fully identicalruns per question over 4 days

What this means for measurement: a single reading of your AI visibility carries meaningful noise. A trend across weekly readings does not. Any tool reporting a precise figure from one run is overstating its own precision.

07 · By question type

Question type changes everything.

The 1,237 questions split into six families. Which sources an engine reaches for depends more on what kind of question it is than on which engine is answering.

Source mix by question family

Where citations come from, per question type

Six families across 1,237 questions. Percentages are of citations within that family.

EditorialReviewsCommunityReferenceVendor
Category discovery
472212
Best X for Y. Editorial roundups dominate. The hardest family to enter and the most valuable to win.
Comparison
393117
X vs Y. Review platforms spike hard here, because engines want a scored, structured comparison.
Pricing
33181525
Highest vendor share of any family. Engines will use your pricing page, if it states numbers plainly.
Integration
212341
Documentation wins outright. Marketing pages are almost never cited on integration questions.
Troubleshooting
164432
Community and docs split it. The highest forum share of any question family by a wide margin.
Brand and trust
343819
Is X any good. Review platforms lead, and your own site is effectively excluded from the answer.
Percentages are of citations within that question family. The spread between families is wider than the spread between engines, which is the more useful finding.

08 · Most-cited domains

The sites AI reads before it answers you.

Every domain cited more than forty times, with which engines used it. Sorted by total citations across the full run.

#DomainTypeEnginesCitationsShare
1reddit.comCommunity51,0425.7%
2g2.comReview platform58264.5%
3wikipedia.orgReference57143.9%
4youtube.comCommunity45833.2%
5capterra.comReview platform44972.7%
6forbes.comMedia44312.3%
7techradar.comIndependent editorial43882.1%
8nerdwallet.comIndependent editorial33441.9%
9trustradius.comReview platform43011.6%
10theverge.comMedia32761.5%
11zapier.comIndependent editorial42641.4%
12pcmag.comIndependent editorial32491.4%
13news.ycombinator.comCommunity32311.3%
14wirecutter.comIndependent editorial32181.2%
15stackoverflow.comReference32041.1%
16trustpilot.comReview platform31871%
17businessinsider.comMedia21630.9%
18softwareadvice.comReview platform31520.8%
19producthunt.comCommunity21410.8%
20gartner.comIndependent editorial21280.7%

Top 20 of 713 cited domains. The full ranked list ships with the dataset.

09 · By category

Your category has its own rules.

The overall split hides real variation. In regulated categories AI leans on institutions; in consumer categories it leans on communities. Where you need to show up depends entirely on which of these you sell into.

Source mix by category

Where citations come from, per category

Eight categories, each with its own retrieval habits. Percentages are of citations within that category.

EditorialReviewsCommunityReferenceVendor
B2B software
38241614
Review platforms are unusually strong. G2 and Capterra alone carry 12% of category citations.
Consumer finance
441227
Regulators and institutional reference sources dominate. The highest reference share of any category.
Ecommerce and retail
462118
Editorial roundups decide almost everything. Wirecutter appears in 14% of answers.
Healthcare
3148
Clinical and government sources crowd out everything else. The hardest category for a brand to enter.
Travel
412617
Review platforms and community threads run neck and neck. TripAdvisor is the single most-cited domain.
Marketplaces
36142913
The highest community share of any category. Seller forums decide the supply-side answers.
Developer tools
292630
Documentation is king. The only category where vendor pages lose to reference sources outright.
Professional services
481715
Editorial and directory listings carry it. Almost no community presence at all.
Percentages are of citations within that category's questions. Rows sum to 100 before rounding.

10 · The Index

The brands AI recommends most.

Share of voice across every answer in the run: how often each brand was named as a proportion of all 2,904 brand mentions. This is the Index itself, and it is a measure of recommendation, not of market share.

Share of all brand mentions

Most-recommended brands across all categories

Of 2,904 brands named across 6,185 answers, top 14 by mention share.

HubSpot
3.4%
Shopify
3.1%
Notion
2.9%
Stripe
2.8%
Slack
2.6%
Salesforce
2.4%
Canva
2.2%
Asana
2.1%
Airtable
1.9%
Figma
1.8%
Zapier
1.7%
Monday.com
1.6%
Intercom
1.5%
Klaviyo
1.4%
No brand exceeded 3.5%. Recommendation is far less concentrated than citation, because engines vary the names they give while returning to the same sources.
3.4%highest single-brand share2,904distinct brands named44%named but never cited2.9brands named per answer

11 · How brands are described

Named is not the same as praised.

Every brand mention was classified by how the surrounding sentence described it. Most mentions are neutral, but the critical ones cluster around a small number of recurring themes.

Tone of brand mentions

How AI describes the brands it names

All 2,904 brand mentions across 6,185 answers, classified by sentence-level tone.

52%29%13%
Neutral, listed only
52%
Positive with a reason
29%
Qualified, with a caveat
13%
Critical
6%
A qualified mention still carries a caveat into the buyer's head. Together, qualified and critical mentions affected 19% of every brand named.

Caveat themes

What the caveats are about

Of every qualified or critical mention, what the reservation was about.

Price or hidden costs
31%
Setup and onboarding time
22%
Support responsiveness
18%
Missing integrations
16%
Reliability or downtime
13%
19%of mentions carry a caveat31%of caveats are about price6%outright critical2.4×caveats from review sites vs news

Nearly every critical mention traced back to a review platform or forum thread rather than to editorial writing. Sentiment is repaired where the reviews live, not on your own site.

12 · Comparison questions

X vs Y questions are won by strangers.

We isolated the 214 head-to-head comparison questions in the set. These are the highest-intent questions a buyer asks, and they are the ones brands lose most decisively.

82%of comparison answers cited neither vendor

The engine went looking for an impartial account and found one somewhere else.

18%cited at least one of the two vendors

Almost always the vendor with a genuinely balanced comparison page, not a sales page.

37%cited a review platform

G2, Capterra and TrustRadius are the default fallback when no neutral article exists.

11%cited a community thread

Reddit outranked both vendors on 46 of the 214 questions.

The vendors that did get cited had one thing in common: a comparison page that named a competitor's advantage in plain language. Pages that only listed their own wins were not cited once.

13 · What changed

Three editions in, the shape is shifting.

Compared against Edition 02, published March 2026, and Edition 01, published November 2025.

Across three editions

What moved between editions

Edition 01 November 2025, Edition 02 March 2026, Edition 03 August 2026.

Vendor-owned citation share139
13E01
11E02
9E03
Falling steadily. Engines are becoming less willing to take a brand's word for it.
Community citation share914
9E01
12E02
14E03
Rising fastest of any source type. Forum archives are being retrieved more, not less.
Citations per answer3.14.3
3.1E01
3.8E02
4.3E03
Answers are getting better sourced across every engine tested.
Top-10 domain concentration3831
38E01
34E02
31E03
Slowly widening. The head is loosening as engines range further.
Answers naming zero brands2212
22E01
17E02
12E03
Engines are increasingly willing to name specific products rather than hedge.
Edition 01 covered three engines, so the citation-per-answer figures are not strictly like for like. Source-share percentages are.

Methodology

How we ran it.

  1. 01
    The question set

    1,237 questions drawn from real buyer language across eight categories, weighted to match the distribution of commercial intent we see in customer prompt sets. The full list ships with the dataset.

  2. 02
    The run

    Every question was submitted to all five engines between 14 and 21 July 2026, from a clean US English context: no account, no personalisation, no conversation history, fresh session per question.

  3. 03
    Capture

    Answers were recorded in full, along with every cited URL and every visible source panel. Where an engine cited the same domain twice in one answer, it was counted once.

  4. 04
    Classification

    Each domain was classified into one of six source types by two reviewers independently, with disagreements resolved by a third. Inter-rater agreement was 94%.

  5. 05
    Brand extraction

    Product and company names were extracted from answer text and normalised against a 12,000-entry entity list. Ambiguous names were reviewed by hand.

  6. 06
    What we did not do

    No paid placements, no logged-in accounts, no regional variation, and no repeated sampling of the same question within the window. Absolute figures will differ from a run in another week; the shape should not.

Cite this research

Use it, quote it, link back to it.

The Prefer Index is published under CC BY 4.0. You are free to reproduce any figure or chart, commercially or otherwise, with attribution.

People also ask

The questions readers ask next, taken from what assistants cluster with this one.

Reddit, G2 and Wikipedia lead, and they are the only three domains cited by all five engines tested. Reddit alone accounted for 5.7% of all 18,412 citations recorded, more than any vendor-owned domain in the entire dataset.

Questions

Asked plainly.

What is The Prefer Index?

A recurring benchmark of how AI answer engines cite sources and recommend brands. Three times a year we run the same 1,237 buyer questions across ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews, record every answer and every citation, and publish the full dataset.

Why does source type matter more than ranking?

Because AI answers are assembled from a handful of retrieved sources rather than ranked from an index. Which kind of site an engine reaches for determines who gets named, so understanding the mix tells you where to invest before it tells you how to optimise.

Is the data available to download?

Yes. The full dataset, all 18,412 citation rows with question, engine, answer text and classified source, is available as a CSV under a permissive licence. Attribution is appreciated, and a link back helps the next edition.

Can I cite this research?

Please do. Cite it as The Prefer Index, Edition 03, Prefer, August 2026, and link to this page. If you are writing about a specific figure, the dataset lets you verify it yourself.

Why do the numbers differ from other studies?

Question set, engine mix and classification method all vary between studies, and answer engines are non-deterministic. Absolute percentages should be treated as approximate. Direction and relative ordering are the durable findings.

How is this different from a keyword ranking study?

A ranking study measures positions in a list of links. This measures which sources an engine chose to build an answer from, and which brands it named as a result. The two correlate loosely and diverge often.

Will you add more engines?

Edition 04 will add Grok and Copilot, and we are evaluating regional engines. The core five will stay constant so the time series remains comparable.

Can I get this run for my own category?

Yes, that is what the product does. A Prefer workspace runs your own prompt set continuously rather than three times a year, sliced by your competitors, markets and buyer personas.

What it means

You do not win AI search on your own website.

Nine in ten citations point somewhere you do not control. That is not a reason to stop publishing, it is a reason to be honest about what publishing does: it makes you quotable once someone else has already put you in the conversation.

01

Earn the editorial

40% of citations sit with independent writers. Getting into the roundups and guides that already rank for your category is the single highest-leverage move available.

02

Fix the review record

16% of citations come from review platforms, and they are the fallback whenever no neutral article exists. Placement, category tags and review volume all read straight into answers.

03

Write the honest comparison

The only vendor pages cited on comparison questions were the ones that named a competitor's advantage. Every purely promotional comparison page in the set was ignored.

04

Show up in the threads

Community share is rising faster than any other source type. The archives being retrieved today were written years ago, which is exactly why starting now matters.

Get your free AI visibility report
in about 10 minutes.

See how answer engines describe your brand today, and where the openings are to outpace the competition.