The AEO Loop: A Weekly AEO Strategy That Wins AI Citations

The AEO loop is Prefer's weekly AEO framework: measure tracked prompts, classify cited sources, read winning pages, score honestly, act. The full recipe.

Javed Khatri Javed Khatri Co-founder, Prefer

16 min read Explainers

Cover for "The AEO Loop: A Weekly AEO Strategy That Wins AI Citations"

The short answer

What is the AEO loop?

The AEO loop is Prefer's weekly process for winning AI citations: measure a fixed portfolio of prompts, classify every source the engines cite, read the winning pages, score your visibility honestly, and ship the fixes the gaps point to. We run it on our own brand every Monday. Run it yourself for about a dollar a week, or let Prefer measure it on five AI engines; Prefer Managed ships the fixes.

Key takeaways

  • The AEO loop is Prefer's five-step weekly process: measure a fixed prompt portfolio, classify every cited source, read the winning pages, score mention and citation rates, then act on the gaps.
  • In Prefer's September 2026 run, half of everything ChatGPT cited in our category was a vendor's own page, and 29% was documentation and reference pages. Owned pages are retrievable answers, not brochures.
  • In our September 2, 2026 ChatGPT run, TechRadar was cited in 12 of the 46 live-search answers that listed sources, across five of its pages. Its review of Ahrefs appeared in 7 of those answers.
  • Only answers that triggered live web search count as citation tests. Short definitional prompts are often answered from the model's memory instead.
  • A 50-prompt portfolio costs about a dollar a week to measure via API. The real cost is the hours spent reading cited pages and shipping fixes, which is also where the results come from.

The AEO loop is Prefer’s weekly answer engine optimization process: measure a fixed portfolio of prompts, classify every source the AI engines cite, read the pages that win, score your visibility honestly, and ship the fixes the gaps point to. Then run it again next Monday.

Most teams treat AI visibility as a report. They buy an audit, read that ChatGPT does not recommend them, feel bad for an afternoon, and move on. The audit was correct the day it ran and stale three weeks later, because AI answers move constantly and the sources behind them move too.

We run the loop on our own brand every week. This post is the full recipe: the prompt portfolio, the source taxonomy, the metrics, the actual costs, and what reading all 134 pages ChatGPT cited in our category taught us about how these engines choose sources. Nothing is held back, because the process only works if you do the work, and if you would rather not do the work, well, that is what we sell.

What is the AEO loop?#

Five steps, one week, repeated. Prefer, our platform, runs the measuring for customers across ChatGPT, Gemini, Perplexity, Google AI Overviews and AI Mode, and Prefer Managed also ships the fixes.

The AEO loop drawn as a cycle of five steps that repeat clockwise every week. 01 Measure: Run the prompt portfolio, record answers and sources. 02 Classify: Sort every cited domain into the seven classes. 03 Read: Read each new cited page, note why it wins. 04 Score: Mention rate, citation rate, share of voice. 05 Act: Ship the top on-site and off-page fixes. In the center: one week per cycle, and each cycle compounds the last.
The five steps of the AEO loop. The output of each week's run, the gap report, becomes the input for the next week's work. The measurement step covers whichever engines you track: ChatGPT, Claude, Gemini, Perplexity, Google AI Overviews.

The first run is just an audit, and a one-time AI visibility audit is the right way in. The loop is what happens when you refuse to let the audit go stale. Week one’s report tells you which prompts you lose and which pages beat you. Week two, you have shipped two fixes and you measure again. The loop turns AI visibility from a report you read into a number you operate on.

Why a loop and not a one-off audit?#

Three reasons, all of them measurement realities rather than marketing preferences.

AI answers are volatile. The same prompt, on the same engine, on the same day, can return a different shortlist. This is the AI equivalent of SERP volatility, except bigger, because a generated answer has more degrees of freedom than ten blue links. One measurement is a snapshot with error bars. A weekly series is a trend you can trust, and the rule we use is simple: a single-week flip is noise until the next run confirms it.

Models change under you. When an engine swaps its default model, your trend line breaks the same way it does after a Google core update. You do not compare numbers across the break, you start a new segment and note the date. If you are not measuring weekly, you will not even know the break happened, and you will draw conclusions from two numbers that were never comparable.

The work compounds. A page you rebuild to answer a prompt keeps answering it. A listicle you earn a slot in keeps getting retrieved. Each week’s fixes stay in the ground while you plant the next ones. Citations do not mechanically become traffic, and we would not claim they do, but recommendation share in the place your buyers ask questions is not a thing you want trending against you.

Step 1: measure a fixed prompt portfolio#

Everything downstream depends on what you measure, so this step gets the most design care.

Build a portfolio of 30 to 50 prompts, phrased the way real buyers actually ask, and spread across four intent tiers:

  • High-intent commercial. “Best [category] for [use case]”, “[your brand] pricing”, “[your brand] vs [competitor]”. The prompts closest to money.
  • Category and competitor. “Best [category] tools in 2026”, “[competitor] alternatives”. Where shortlists form.
  • Vertical and persona. “[Category] for B2B SaaS”, “for small agencies”. Lower volume, high fit.
  • Informational. “How do I [problem your product solves]?” Where buyers start.

Split branded prompts from non-branded ones and never blend the two in a headline number. Being mentioned when someone asks about you by name is table stakes. Being mentioned when someone asks “what is the best X” is the game.

Two rules keep the data honest over months:

  1. A prompt’s text is frozen once you start tracking it. Rewording changes what you are measuring. If a prompt needs a rewrite, retire it and add a new one, the way you would never silently edit a survey question mid-study.
  2. Pin the model and record it. We measure against gpt-5.6-luna because it is what free-tier ChatGPT users get. When the pin changes, the trend segment ends.

For the instrument, we use DataForSEO’s LLM Responses endpoint: it runs a real ChatGPT query with live web search and returns the full answer, every cited URL, and the list of searches the engine ran to build the answer. Our last 50-prompt run metered $0.95, about two cents per prompt on the standard queue. You could also run prompts by hand in a browser every Monday. It works, it just does not scale past a dozen prompts, and hand-copying source lists gets old by week two.

Whichever instrument you use, record four things per prompt: the answer text, every brand mentioned and in what order, every source cited, and the engine’s fan-out searches, the queries it ran behind the scenes.

Anatomy of one tracked prompt, illustrative. Prompt: What are the best project tracking tools for small teams? The answer recommends, in order: 1 TaskFlow, 2 Acme, 3 Boardly, 4 Trackette; Acme is mentioned second. Cited sources: techreview.example (editorial, “I tested 7 project trackers” listicle); taskflow.example (vendor-owned, TaskFlow's own comparison page); reddit.com (community, “What do you actually use?” thread). Acme's own site is not cited. The engine's fan-out searches: best project tracking tools 2026; site:taskflow.example pricing; Acme project tracker reviews. You record: mentioned yes at position 2, cited no.
An illustrative tracked-prompt result (fictional brands). One prompt returns four recordable facts: who the answer recommends and in what order, which pages it cites, which searches the engine ran to build it, and whether live web search fired at all.

One flag matters more than any other: did the engine actually search the web, or did it answer from memory? Long, specific prompts usually trigger live search. Short definitional ones often do not, and an answer assembled from training data is not a citation test, it is a memory test. In our September 2 run, 47 of 50 prompts triggered live search. The other 3 were answered from memory and excluded from scoring: two short definitions (“What is answer engine optimization (AEO)?” and “What is the difference between AEO and SEO?”) and “Profound alternatives”, which the engine read as a request for synonyms of the word “profound”. Whatever tooling you use, capture this flag and score only the live-search answers.

Step 2: classify every source the engines cite#

The list of cited URLs is where the loop stops being a scoreboard and starts being a strategy. Every domain the engines cite falls into one of seven classes:

ClassWhat it isWhat you can do about it
Vendor-ownedA vendor’s own site (yours or a competitor’s)Ship answer-shaped pages (the step 3 pattern)
Independent editorialThird-party listicles, press, newslettersPitch for inclusion
CommunityReddit, YouTube, forums, HNParticipate authentically
Review profilesG2, Capterra, directoriesComplete and maintain your listings
Documentation and referenceOfficial docs, Wikipedia, papersRarely actionable, occasionally citable-adjacent
SeededCoordinated look-alike ranking sitesFlag, discount, never imitate
UnknownNot yet classifiedRead it this week

Here is what that mix actually looked like when we classified all 134 unique URLs ChatGPT cited across our own September portfolio run:

Bar chart. Share of the 134 unique URLs ChatGPT cited across one 50-prompt category run: Vendor-owned pages 50.0%, Docs and reference 29.1%, Independent editorial 12.7%, Review-site profiles 5.2%, Community threads 2.2%, Seeded or coordinated 0.7%.
Source classes of the 134 unique URLs cited in Prefer's 2026-09-02 tracked-prompt run (50 prompts, ChatGPT with live web search, AI-visibility category). Vendor-owned pages, ours and competitors' combined, were half of everything cited. Category mixes differ; measure your own. Related: our Reddit citation study
Reuse this chart

Free to reuse with attribution. Copy the embed code, or save the image.

Download PNG
<figure>
  <img src="https://tryprefer.com/figures/aeo-loop-source-classes.png" alt="Bar chart. Share of the 134 unique URLs ChatGPT cited across one 50-prompt category run: Vendor-owned pages 50.0%, Docs and reference 29.1%, Independent editorial 12.7%, Review-site profiles 5.2%, Community threads 2.2%, Seeded or coordinated 0.7%." width="868" height="544" style="max-width:100%;height:auto" />
  <figcaption>What ChatGPT cites: half of it is vendors' own pages by <a href="https://tryprefer.com/blog/aeo-loop/">Prefer</a></figcaption>
</figure>

Two things in that chart surprised us. First, half of everything ChatGPT cited was a vendor’s own page: pricing pages, comparison pages, docs. The lazy read of AI search is “third parties decide everything now.” The data says otherwise. Engines retrieve vendor pages constantly, when those pages are shaped like answers. Your site is not a brochure in this game, it is a source.

Second, documentation and reference was 29.1%, and almost all the how-to prompts in the run were answered from official docs with no brands cited at all. Independent editorial, the listicles everyone obsesses over, was 12.7%. Review profiles 5.2%, community threads 2.2%, seeded sites 0.7%. Your category’s mix will differ (in our study of software buyer prompts, Reddit alone was 30% of ChatGPT’s citations, a very different shape), which is exactly why you classify your own data instead of borrowing someone else’s chart. And this is one run, in one category, measured on September 2, 2026: treat the exact shares as directional, the way you would any single crawl.

One class deserves a warning. Seeded sites are the AI era’s private blog network: fresh domains full of thin “best tools in 2026” listicles that exist to feed one vendor into AI answers. We watched one vendor in our category surge to the top of four different prompt types on the strength of a coordinated ring of look-alike ranking sites. The tells: a domain you have never heard of, registered recently, no named authors, and every list crowning the same product. Flag these, exclude them from your competitive read, and never copy the tactic. Right now the engines still retrieve this stuff, which is why it works. They will get better at discounting it, the same way Google eventually got better at link networks, and you do not want your brand in the blast radius when they do.

Step 3: read every page the engines cite#

This is the step everyone skips, and it is where the actual strategy comes from. A citation count tells you that you are losing. Only the pages themselves tell you why.

Read every newly cited URL, every week, and answer four questions about each:

  1. What is it? Listicle, comparison, docs, community thread, review profile, data study?
  2. Why does the engine cite it? A direct answer in the first sentence? Comparison tables? Named numbers and dates? Fresh update stamp? Schema?
  3. Who does it recommend? Which brands, in what order, and is yours there?
  4. How would you own it better? Beat it with a better page, join it with an honest contribution, pitch it, or fix your listing?

After reading 134 of them in one week, the pattern behind winning pages was boringly consistent: a direct answer in the first sentence, a comparison table with dated pricing, original data with a stated methodology, and an honest concession about where rivals win. That last one surprised us most. Several of the cited pages admitted where rivals win, so being honest about weak spots did not keep them out of the answers.

The reading also protects you from a mistake the counts alone would cause. Some of the “editorial” pages cited in our category turned out to be vendor-controlled on inspection: a product hub whose number-one slot is a paid placement, and a trade publication owned by a competitor. If we had pitched those as neutral press, we would have wasted weeks. Classify first, then act.

Step 4: score it honestly#

Four numbers, computed the same way every week. No adjustments, no charitable rounding.

  • Mention rate. On how many live-search answers is your brand named? Track branded and non-branded separately.
  • Citation rate. On how many answers does your own domain appear as a source? This is a different number from mention rate: you can be recommended without being the source, and cited inside a competitor’s recommendation.
  • Share of voice. For every brand in your category, how often does it appear and at what position? This is where you learn who actually owns your prompts.
  • The fan-out check. Does the engine ever search for your brand by name when building an answer? Engines run background queries like “site:vendor.com pricing” for brands they already trust. If your brand never appears in any fan-out, you have an entity problem, not a content problem, and the fix starts with profiles, consistent facts, and third-party references rather than another blog post.

Here is what a weekly scoreboard looks like, with fictional numbers for a fictional brand:

Metric (Acme, week 12)This weekLast week
Live-search answers36 of 40 prompts35 of 40
Mentioned (non-branded)9 of 368 of 35
Cited (acme.example in sources)4 of 364 of 35
Share of voice position#4 in category#5

Two scoring rules keep you honest. A single-week flip, up or down, is noise until the next run confirms it, so Acme’s marketing team does not get to celebrate the ninth mention yet. And when the engine’s default model changes, the series breaks: label the segments and never compare across them, just as you would not compare rankings across a core update as if nothing happened.

Step 5: act, then run it again#

The report ends with a queue, or the whole exercise was a spectator sport. We split ours in two.

On-site. For each losing prompt worth winning, write a one-page citation brief: the prompt, who wins it, the cited pages powering them, the pattern those pages share, and the spec for your page, including what you can add that the winners lack. Original data, real screenshots, honest pricing. Then build or rebuild the page. Prioritize by intent tier: losing a “best [category]” prompt costs more than losing a how-to.

Off-site. Group the earned targets by class, because each class has exactly one legitimate move. Community threads that engines keep citing: contribute a genuinely useful answer, with your affiliation disclosed. Editorial listicles: pitch the author, having actually read their inclusion criteria. Review profiles: complete them, they get cited for pricing and alternatives more than most brands realize.

Ship two or three actions a week. Not ten. The loop’s power is that next Monday’s run tells you whether last week’s fixes moved anything, and a short queue you actually ship beats a long one you admire.

What the loop taught us about AI citations#

Four things we believed differently before we started measuring every week:

Vendor-owned pages are half the game. We assumed AI answers were built mostly from third-party consensus. In our category run, 50% of cited URLs were vendors’ own pages. The brands winning citations have pages shaped like answers: dated pricing tables, direct first-sentence claims, honest comparisons.

One site showed up in a fifth to a quarter of our answers. In our September 2, 2026 ChatGPT run, TechRadar was cited in 12 of the 46 live-search answers that listed sources, across five of its pages, and its review of Ahrefs alone appeared in 7. Five days later, on September 7, it was cited in 10 of the 47 that listed sources, across six pages. Earned coverage is not a vanity metric in AI search; one strong review on the right site can show up across many prompts.

The how-to prompts are unowned. Documentation and reference pages (platform help centers, developer docs and arXiv papers) were 29.1% of the unique URLs cited, and the how-to prompts were answered from docs with no brands present at all. Whoever writes the genuinely best third-party guide for those questions is walking into empty territory.

Prompt phrasing decides whether you are even measuring. Short prompts, especially definitions, can skip live search and answer from memory. If your tracking is built on them, you are benchmarking the model’s training data, not your visibility. Phrase tracked prompts the way buyers talk, in full questions.

How to run the AEO loop yourself#

The whole thing, one week, about half a day of desk work plus whatever you ship:

One week in the AEO loop, five steps with time and output. Pull the portfolio (20 min, ~$1 in API calls): A dated run file: every answer, mention, source, and fan-out search. Classify new domains (30 min): Every cited domain sorted into the seven source classes. Read new cited pages (2-3 hrs): One short note per page: why the engine picked it. Score and report (1 hr): Mention rate, citation rate, share of voice, and the week's gap list. Act on the top gaps (rest of week): 2-3 shipped fixes: a rebuilt page, a pitch, an honest community answer.
A working week in the AEO loop for a 50-prompt portfolio. Times are what the steps take us in practice; the API cost is metered DataForSEO LLM Responses usage, and our last 50-prompt run cost $0.95. The acting step is bounded by choice: two or three shipped fixes beat a ten-item queue.

Your first-week checklist:

  1. Write 30 to 50 prompts across the four tiers, branded split from non-branded.
  2. Freeze the prompt texts. Note the engine and model you are measuring.
  3. Run the portfolio: by hand in a browser, or about a dollar through an API.
  4. Record mentions, order, sources, and fan-out searches for every prompt. Discard answers that did not trigger live search.
  5. Classify every cited domain into the seven classes.
  6. Read every cited page and answer the four questions.
  7. Compute mention rate, citation rate, and share of voice. Write down the top three gaps and ship against them.

Then do it again next Monday, and diff.

There are three honest ways to run this. Manually, free, fine up to a dozen prompts. Instrumented yourself, about a dollar a week plus half a day of reading and classifying. Or have us run it for you: the audit that starts the loop is free, and the platform is the loop with the manual labor automated, plus a team shipping the fixes on Prefer Managed. Whichever you pick, the loop is the same. Measure, classify, read, score, act. The brands that get recommended in AI answers next year are the ones running it this year.

People also ask

Frequently asked questions.

Updated 19 September 2026

What is the AEO loop?

The AEO loop is a weekly answer engine optimization process with five steps: measure a fixed portfolio of tracked prompts against the AI engines, classify every source the answers cite, read the pages that win citations, score your brand's mention rate and citation rate, and act on the biggest gaps. Prefer runs it on its own brand every week. For customers it runs the measuring across five AI engines and writes content, and Prefer Managed also ships the off-site fixes. Each week's report sets up the next week's work, which is what makes it a loop rather than an audit.

How many prompts do you need to track?

30 to 50 is enough for one brand; Prefer's Starter plan tracks 30 and Growth tracks 150. Spread them across four intent tiers: high-intent commercial prompts, category and competitor prompts, vertical or persona prompts, and informational prompts. Split branded from non-branded so you never mistake people asking about you for the engines recommending you. Keep each prompt's text frozen once you start tracking it, otherwise your trend line breaks.

How often should you run the AEO loop?

Weekly, if you run it yourself; Prefer handles the measuring step for you on every plan. AI answers are volatile, the same prompt can return a different answer on a different day, so a single measurement is a snapshot with error bars. A weekly cadence smooths the noise, catches model updates early, and keeps the action queue short enough to actually ship. A one-off audit is the right entry point, but it goes stale in weeks.

How much does the AEO loop cost to run?

About a dollar a week in API calls if you run it yourself, or $24 a month to have Prefer run the measuring for you (30 tracked prompts on five AI engines, plus 15 AI articles). For a 50-prompt portfolio we use DataForSEO's LLM Responses endpoint, which runs a real ChatGPT query with live web search for roughly two cents per prompt on the standard queue; our last 50-prompt run metered $0.95. The bigger cost is time: plan on half a day a week for classifying sources, reading newly cited pages, and writing up the gaps, plus whatever it takes to ship the fixes.

Is the AEO loop different for GEO?

No. GEO (generative engine optimization) and AEO (answer engine optimization) are two names for the same discipline, and Prefer runs the same loop for both. You measure prompts, classify sources, read winners, score, and act, whatever you call the practice. The engines do not care which acronym you use.

Get your free AI visibility report
in about 10 minutes.

See how answer engines describe your brand today, and where the openings are to outpace the competition.