The AEO loop is Prefer’s weekly answer engine optimization process: measure a fixed portfolio of prompts, classify every source the AI engines cite, read the pages that win, score your visibility honestly, and ship the fixes the gaps point to. Then run it again next Monday.
Most teams treat AI visibility as a report. They buy an audit, read that ChatGPT does not recommend them, feel bad for an afternoon, and move on. The audit was correct the day it ran and stale three weeks later, because AI answers move constantly and the sources behind them move too.
We run the loop on our own brand every week. This post is the full recipe: the prompt portfolio, the source taxonomy, the metrics, the actual costs, and what reading all 134 pages ChatGPT cited in our category taught us about how these engines choose sources. Nothing is held back, because the process only works if you do the work, and if you would rather not do the work, well, that is what we sell.
What is the AEO loop?#
Five steps, one week, repeated. Prefer, our platform, runs the measuring for customers across ChatGPT, Gemini, Perplexity, Google AI Overviews and AI Mode, and Prefer Managed also ships the fixes.
The first run is just an audit, and a one-time AI visibility audit is the right way in. The loop is what happens when you refuse to let the audit go stale. Week one’s report tells you which prompts you lose and which pages beat you. Week two, you have shipped two fixes and you measure again. The loop turns AI visibility from a report you read into a number you operate on.
Why a loop and not a one-off audit?#
Three reasons, all of them measurement realities rather than marketing preferences.
AI answers are volatile. The same prompt, on the same engine, on the same day, can return a different shortlist. This is the AI equivalent of SERP volatility, except bigger, because a generated answer has more degrees of freedom than ten blue links. One measurement is a snapshot with error bars. A weekly series is a trend you can trust, and the rule we use is simple: a single-week flip is noise until the next run confirms it.
Models change under you. When an engine swaps its default model, your trend line breaks the same way it does after a Google core update. You do not compare numbers across the break, you start a new segment and note the date. If you are not measuring weekly, you will not even know the break happened, and you will draw conclusions from two numbers that were never comparable.
The work compounds. A page you rebuild to answer a prompt keeps answering it. A listicle you earn a slot in keeps getting retrieved. Each week’s fixes stay in the ground while you plant the next ones. Citations do not mechanically become traffic, and we would not claim they do, but recommendation share in the place your buyers ask questions is not a thing you want trending against you.
Step 1: measure a fixed prompt portfolio#
Everything downstream depends on what you measure, so this step gets the most design care.
Build a portfolio of 30 to 50 prompts, phrased the way real buyers actually ask, and spread across four intent tiers:
- High-intent commercial. “Best [category] for [use case]”, “[your brand] pricing”, “[your brand] vs [competitor]”. The prompts closest to money.
- Category and competitor. “Best [category] tools in 2026”, “[competitor] alternatives”. Where shortlists form.
- Vertical and persona. “[Category] for B2B SaaS”, “for small agencies”. Lower volume, high fit.
- Informational. “How do I [problem your product solves]?” Where buyers start.
Split branded prompts from non-branded ones and never blend the two in a headline number. Being mentioned when someone asks about you by name is table stakes. Being mentioned when someone asks “what is the best X” is the game.
Two rules keep the data honest over months:
- A prompt’s text is frozen once you start tracking it. Rewording changes what you are measuring. If a prompt needs a rewrite, retire it and add a new one, the way you would never silently edit a survey question mid-study.
- Pin the model and record it. We measure against gpt-5.6-luna because it is what free-tier ChatGPT users get. When the pin changes, the trend segment ends.
For the instrument, we use DataForSEO’s LLM Responses endpoint: it runs a real ChatGPT query with live web search and returns the full answer, every cited URL, and the list of searches the engine ran to build the answer. Our last 50-prompt run metered $0.95, about two cents per prompt on the standard queue. You could also run prompts by hand in a browser every Monday. It works, it just does not scale past a dozen prompts, and hand-copying source lists gets old by week two.
Whichever instrument you use, record four things per prompt: the answer text, every brand mentioned and in what order, every source cited, and the engine’s fan-out searches, the queries it ran behind the scenes.
One flag matters more than any other: did the engine actually search the web, or did it answer from memory? Long, specific prompts usually trigger live search. Short definitional ones often do not, and an answer assembled from training data is not a citation test, it is a memory test. In our September 2 run, 47 of 50 prompts triggered live search. The other 3 were answered from memory and excluded from scoring: two short definitions (“What is answer engine optimization (AEO)?” and “What is the difference between AEO and SEO?”) and “Profound alternatives”, which the engine read as a request for synonyms of the word “profound”. Whatever tooling you use, capture this flag and score only the live-search answers.
Step 2: classify every source the engines cite#
The list of cited URLs is where the loop stops being a scoreboard and starts being a strategy. Every domain the engines cite falls into one of seven classes:
| Class | What it is | What you can do about it |
|---|---|---|
| Vendor-owned | A vendor’s own site (yours or a competitor’s) | Ship answer-shaped pages (the step 3 pattern) |
| Independent editorial | Third-party listicles, press, newsletters | Pitch for inclusion |
| Community | Reddit, YouTube, forums, HN | Participate authentically |
| Review profiles | G2, Capterra, directories | Complete and maintain your listings |
| Documentation and reference | Official docs, Wikipedia, papers | Rarely actionable, occasionally citable-adjacent |
| Seeded | Coordinated look-alike ranking sites | Flag, discount, never imitate |
| Unknown | Not yet classified | Read it this week |
Here is what that mix actually looked like when we classified all 134 unique URLs ChatGPT cited across our own September portfolio run:
Two things in that chart surprised us. First, half of everything ChatGPT cited was a vendor’s own page: pricing pages, comparison pages, docs. The lazy read of AI search is “third parties decide everything now.” The data says otherwise. Engines retrieve vendor pages constantly, when those pages are shaped like answers. Your site is not a brochure in this game, it is a source.
Second, documentation and reference was 29.1%, and almost all the how-to prompts in the run were answered from official docs with no brands cited at all. Independent editorial, the listicles everyone obsesses over, was 12.7%. Review profiles 5.2%, community threads 2.2%, seeded sites 0.7%. Your category’s mix will differ (in our study of software buyer prompts, Reddit alone was 30% of ChatGPT’s citations, a very different shape), which is exactly why you classify your own data instead of borrowing someone else’s chart. And this is one run, in one category, measured on September 2, 2026: treat the exact shares as directional, the way you would any single crawl.
One class deserves a warning. Seeded sites are the AI era’s private blog network: fresh domains full of thin “best tools in 2026” listicles that exist to feed one vendor into AI answers. We watched one vendor in our category surge to the top of four different prompt types on the strength of a coordinated ring of look-alike ranking sites. The tells: a domain you have never heard of, registered recently, no named authors, and every list crowning the same product. Flag these, exclude them from your competitive read, and never copy the tactic. Right now the engines still retrieve this stuff, which is why it works. They will get better at discounting it, the same way Google eventually got better at link networks, and you do not want your brand in the blast radius when they do.
Step 3: read every page the engines cite#
This is the step everyone skips, and it is where the actual strategy comes from. A citation count tells you that you are losing. Only the pages themselves tell you why.
Read every newly cited URL, every week, and answer four questions about each:
- What is it? Listicle, comparison, docs, community thread, review profile, data study?
- Why does the engine cite it? A direct answer in the first sentence? Comparison tables? Named numbers and dates? Fresh update stamp? Schema?
- Who does it recommend? Which brands, in what order, and is yours there?
- How would you own it better? Beat it with a better page, join it with an honest contribution, pitch it, or fix your listing?
After reading 134 of them in one week, the pattern behind winning pages was boringly consistent: a direct answer in the first sentence, a comparison table with dated pricing, original data with a stated methodology, and an honest concession about where rivals win. That last one surprised us most. Several of the cited pages admitted where rivals win, so being honest about weak spots did not keep them out of the answers.
The reading also protects you from a mistake the counts alone would cause. Some of the “editorial” pages cited in our category turned out to be vendor-controlled on inspection: a product hub whose number-one slot is a paid placement, and a trade publication owned by a competitor. If we had pitched those as neutral press, we would have wasted weeks. Classify first, then act.
Step 4: score it honestly#
Four numbers, computed the same way every week. No adjustments, no charitable rounding.
- Mention rate. On how many live-search answers is your brand named? Track branded and non-branded separately.
- Citation rate. On how many answers does your own domain appear as a source? This is a different number from mention rate: you can be recommended without being the source, and cited inside a competitor’s recommendation.
- Share of voice. For every brand in your category, how often does it appear and at what position? This is where you learn who actually owns your prompts.
- The fan-out check. Does the engine ever search for your brand by name when building an answer? Engines run background queries like “site:vendor.com pricing” for brands they already trust. If your brand never appears in any fan-out, you have an entity problem, not a content problem, and the fix starts with profiles, consistent facts, and third-party references rather than another blog post.
Here is what a weekly scoreboard looks like, with fictional numbers for a fictional brand:
| Metric (Acme, week 12) | This week | Last week |
|---|---|---|
| Live-search answers | 36 of 40 prompts | 35 of 40 |
| Mentioned (non-branded) | 9 of 36 | 8 of 35 |
| Cited (acme.example in sources) | 4 of 36 | 4 of 35 |
| Share of voice position | #4 in category | #5 |
Two scoring rules keep you honest. A single-week flip, up or down, is noise until the next run confirms it, so Acme’s marketing team does not get to celebrate the ninth mention yet. And when the engine’s default model changes, the series breaks: label the segments and never compare across them, just as you would not compare rankings across a core update as if nothing happened.
Step 5: act, then run it again#
The report ends with a queue, or the whole exercise was a spectator sport. We split ours in two.
On-site. For each losing prompt worth winning, write a one-page citation brief: the prompt, who wins it, the cited pages powering them, the pattern those pages share, and the spec for your page, including what you can add that the winners lack. Original data, real screenshots, honest pricing. Then build or rebuild the page. Prioritize by intent tier: losing a “best [category]” prompt costs more than losing a how-to.
Off-site. Group the earned targets by class, because each class has exactly one legitimate move. Community threads that engines keep citing: contribute a genuinely useful answer, with your affiliation disclosed. Editorial listicles: pitch the author, having actually read their inclusion criteria. Review profiles: complete them, they get cited for pricing and alternatives more than most brands realize.
Ship two or three actions a week. Not ten. The loop’s power is that next Monday’s run tells you whether last week’s fixes moved anything, and a short queue you actually ship beats a long one you admire.
What the loop taught us about AI citations#
Four things we believed differently before we started measuring every week:
Vendor-owned pages are half the game. We assumed AI answers were built mostly from third-party consensus. In our category run, 50% of cited URLs were vendors’ own pages. The brands winning citations have pages shaped like answers: dated pricing tables, direct first-sentence claims, honest comparisons.
One site showed up in a fifth to a quarter of our answers. In our September 2, 2026 ChatGPT run, TechRadar was cited in 12 of the 46 live-search answers that listed sources, across five of its pages, and its review of Ahrefs alone appeared in 7. Five days later, on September 7, it was cited in 10 of the 47 that listed sources, across six pages. Earned coverage is not a vanity metric in AI search; one strong review on the right site can show up across many prompts.
The how-to prompts are unowned. Documentation and reference pages (platform help centers, developer docs and arXiv papers) were 29.1% of the unique URLs cited, and the how-to prompts were answered from docs with no brands present at all. Whoever writes the genuinely best third-party guide for those questions is walking into empty territory.
Prompt phrasing decides whether you are even measuring. Short prompts, especially definitions, can skip live search and answer from memory. If your tracking is built on them, you are benchmarking the model’s training data, not your visibility. Phrase tracked prompts the way buyers talk, in full questions.
How to run the AEO loop yourself#
The whole thing, one week, about half a day of desk work plus whatever you ship:
Your first-week checklist:
- Write 30 to 50 prompts across the four tiers, branded split from non-branded.
- Freeze the prompt texts. Note the engine and model you are measuring.
- Run the portfolio: by hand in a browser, or about a dollar through an API.
- Record mentions, order, sources, and fan-out searches for every prompt. Discard answers that did not trigger live search.
- Classify every cited domain into the seven classes.
- Read every cited page and answer the four questions.
- Compute mention rate, citation rate, and share of voice. Write down the top three gaps and ship against them.
Then do it again next Monday, and diff.
There are three honest ways to run this. Manually, free, fine up to a dozen prompts. Instrumented yourself, about a dollar a week plus half a day of reading and classifying. Or have us run it for you: the audit that starts the loop is free, and the platform is the loop with the manual labor automated, plus a team shipping the fixes on Prefer Managed. Whichever you pick, the loop is the same. Measure, classify, read, score, act. The brands that get recommended in AI answers next year are the ones running it this year.
People also ask
- What is the AEO loop?
- How do you measure AI visibility?
- How often should you check your brand's AI visibility?
- How much does it cost to track AI citations?