You cannot improve a number you have never measured, and almost nobody has measured this one. Most teams have a screenshot of one ChatGPT answer and a gut feeling. The audit replaces that with a real baseline: the questions your buyers ask AI, how often you show up across engines today, and what the models actually say about you. It is the first move in the plan, and every later chapter is judged against the number you set here.
Two ways to do it. By hand, with the engines and a spreadsheet, in about an hour: that is Plays 04 to 07 below. Or run the free AI visibility checker, which does the same across ChatGPT, Claude, and Perplexity in about fifteen minutes and keeps the verbatim answers. Either way, you leave this chapter with a baseline scorecard and your ten money prompts.
Start with the questions, not your brand#
The whole audit hangs on one input: the ten questions you measure against. Get these right and everything downstream (your baseline, your gap list, your whole plan) is aimed at what buyers actually ask. Get them wrong (especially by putting your brand name in them) and you measure a flattering illusion.
Write your 10 money prompts
List the ten buyer questions your category is decided on, so everything you measure and optimize is aimed at what buyers actually ask.
Buyers ask assistants full, natural questions, and our study used exactly those: 47 real buyer questions across 8 industries. Measuring against the questions that decide purchases (not your brand name) is the only baseline worth having. (Prefer AI Citation Study, checked 2026-07-10)
- Write 5 'best [category] for [buyer]' prompts, one per buyer segment you sell to.
- Write 3 '[you] vs [rival]' or '[rival] vs [rival]' prompts for your closest head-to-heads.
- Write 2 'how do I choose a [category]' advice prompts.
- Strip your brand name out of the 'best' and 'how' prompts; keep it only in the head-to-heads. Test the category question, not a vanity question.
- Freeze the set. These ten, unchanged, are what you re-run every week; a moving set cannot show a trend.
Done when: Ten fixed prompts written, brand-neutral where it counts.
Verify it worked: Read each as a skeptical buyer. Would they actually type this into ChatGPT? If not, rewrite it.
Common failure mode: Branded prompts like 'is Acme good?'. The model will always describe you when you name you, so branded prompts inflate your score and hide the real gap.
Run the baseline#
Now measure. Ask the ten prompts, and score each answer one of three ways. This is the distinction that makes the number honest: being named is not the same as being cited.
Run your baseline citation check
Get your starting number: how often AI names you across engines today, scored honestly.
The influence happens inside answers you cannot see in analytics: Pew found only 1% of users click the links inside an AI summary. If you do not ask the questions and log the answers, you are flying blind on the channel that increasingly decides shortlists. (Pew Research, July 2025, checked 2026-07-14)
Before you start: Your 10 money prompts from Play 04.
- Run each of the 10 prompts through ChatGPT and at least one grounded engine (Perplexity cites the most), with web search on.
- Confirm the answer actually searched; if it replied from memory with no sources, rephrase to force live search (see chapter 2).
- Score each answer per engine: cited (named and sourced), mentioned (named, no source), or absent.
- Save the verbatim answers; they are your evidence and your Play 06 input.
- Compute your citation share: answers naming you, divided by prompts, per engine and blended.
Done when: Every prompt scored per engine, and your citation share computed.
Verify it worked: Have a colleague run the same ten prompts. Your scores should agree within a point or two; big gaps mean the prompts are ambiguous.
Common failure mode: Scoring memory answers. A no-search answer measures the model's training data, not your live visibility; force grounding first.
Read the answers, not just the score
Audit what AI actually says about you, because a wrong or outdated answer costs you even when it names you.
The answer is the product the buyer acts on. A model that names you but describes you wrong, or with year-old facts, sends a worse signal than a clean absence. We have watched ChatGPT confidently describe the wrong company when asked about our own brand, which no citation count would have caught.
Before you start: The verbatim answers you saved in Play 05.
- For every answer that names you, check three things: is the description accurate, is it current, and is it complete?
- For every answer that does not name you, write down which competitors it names. That is your gap, in brand terms.
- Flag any hallucination or stale fact (wrong pricing, wrong category, a feature you dropped or added).
- Save the worst examples; they are the fastest wins and the clearest proof of why this work matters.
Done when: Each answer about you is tagged accurate, outdated, or wrong, and the competitors named instead of you are listed.
Verify it worked: Ask: could a buyer act correctly on what AI just said about you? If not, you have found a fix, not a mention.
Common failure mode: Celebrating a mention that is actually describing you wrong. A confident, incorrect answer is a problem, not a win.
Set your scorecard and a realistic target
Record your baseline and set a day-90 target in the indicators that actually move first, so you judge progress honestly.
A baseline with no target is trivia, and a target on the wrong metric quits a working plan early. Citation share lags by weeks (see the 90-day chapter), so the day-90 target belongs on leading indicators: placements landed, reviews gained, and gap-list domains that start mentioning you. (The 90-day plan)
- Record your citation share per engine and blended, with today's date. This is the line everything is measured against.
- Note the competitors named instead of you (from Play 06) as your starting gap.
- Set a day-90 target in leading indicators, not citation share: for example, 15 directories live, 10 reviews, and 5 gap-list domains now mentioning you.
- Put the scorecard where you will re-run it weekly (this becomes the scoreboard in Play 32).
Done when: A dated baseline and a leading-indicator target are written down.
Verify it worked: Re-run the ten prompts weekly. The trend across runs, not any single run, is your real measure.
Common failure mode: Setting a citation-share target for day 30. It lags, so you will read a working plan as a failure and stop right before it pays off.
Where this is easy to get wrong#
The audit’s most common failure is measuring the wrong thing and trusting it. Our first baseline did exactly that. We ran too few prompts, some with our brand name in them, on a single pass, and the number bounced around and flattered us. When we fixed it (ten fixed, mostly brand-neutral prompts, run consistently, scored cited-versus-mentioned-versus-absent), the real number was lower and far more useful, because now it could actually move and we could see it move. A baseline you can game is a baseline that lies to you. Strip the brand names, fix the prompt set, and read the trend, not the first run.
Your baseline scorecard + 10 money prompts
Fill this in as you run the plays. Your money prompts and baseline feed every later chapter, and it all saves locally as you type.
Everything you type saves in this browser and assembles into one document on the Your plan page, where you can copy or download it. Nothing is sent anywhere. A duplicatable Notion and Google Sheets version ships with the companion pack.
Where this goes next#
Your baseline and your ten money prompts are the inputs the rest of the playbook runs on. The prompts drive the category classification and the weekly scoreboard; the baseline is the line your 90-day plan is judged against; and the competitors and gaps you found here become your off-page target list. Run the audit today, by hand or with the free checker; to keep it running, Prefer tracks your AI citations across engines inside answer engine insights. Everything you fill in above lands on your plan page.
People also ask
- How do I check if ChatGPT mentions my brand?
- What is a good AI citation rate?
- How do I measure AI search visibility?
- What is share of voice in AI search?