You track brand mentions in AI answers by running a fixed set of buyer prompts on a schedule, always on the same model, and recording whether your brand was named, in what position, and which sources the answer used. Start in a spreadsheet, one row per prompt per week, and move to a tool only once the runs get too big to do by hand. Before any of that, get one distinction straight: a mention is your name in the answer, and a citation is your URL in the sources. They are not the same, and measuring only one hides half the picture.
This guide covers the difference that trips teams up, the manual method step by step, the spreadsheet columns to use, an honest look at a real baseline, and the point where a tool starts to earn its keep.
What is the difference between a mention and a citation?#
A mention is your brand named in the answer text; a citation is your URL listed as a source the answer relied on. The two come apart constantly. An AI can recommend your product by name while its sources are all review sites and Reddit threads, so you have a mention with no citation. It can also pull a fact from your page and list your URL as a source without ever naming your product in the prose, so you have a citation with no mention. Prefer records both for every tracked prompt, so you can see which one you are missing.
Both matter, and they answer different questions. A mention tells you the model believes you belong in the answer. A citation tells you the model read your page and trusted it as evidence. If you track only clicks or only citations, you miss the recommendations that carry no link, which are often the ones that actually move a buyer.
For the precise definitions, see citation rate and brand mention monitoring. The short version is that you want to record both columns for every prompt you run.
How do you track AI brand mentions manually?#
Run a fixed set of real buyer prompts, always on the same model, in a fresh session each time, and log the mention, the position, and the sources for every answer. The method is simple. The discipline is in keeping the prompts and the model fixed so this week compares to last week.
Here is the sequence in practice.
- Write 10 to 50 real buyer prompts. Use the actual questions people ask, grouped by family: category (“best tools for X”), competitor (“X alternatives”), pricing, and use case. Fixed wording keeps runs comparable.
- Pick one engine and model, and note the version. Run every prompt on that same model each week. Switching models mid-track breaks the comparison.
- Ask each prompt in a fresh session. Turn off memory and history so the answer reflects the model, not your past chats.
- Record the result. For every prompt, log whether your brand was mentioned, its position in any list, whether your URL was cited, and which sources the answer used.
- Compute share of voice. Count the prompts that named you versus each competitor, and turn it into a percentage you can watch move.
What columns should your tracking spreadsheet have?#
One row per prompt per week, with columns for date, prompt, engine, model, mention, position, citation, and the sources listed. A spreadsheet is the honest way to start, and it stays useful up to a few dozen prompts. Use these columns:
| Column | What to record |
|---|---|
| Date | The day you ran the prompt (fixed weekly day) |
| Prompt | The exact question, worded identically each week |
| Prompt family | Category, competitor, pricing, or use case |
| Engine | ChatGPT, Perplexity, Gemini, Copilot, or Claude |
| Model | The model version, so runs stay comparable |
| Mentioned | Yes or no: was your brand named in the answer |
| Position | Where you appeared in any ranked list (1, 2, 3…) |
| Cited | Yes or no: was your URL in the sources |
| Sources listed | The domains the answer cited (Reddit, G2, competitors) |
| Competitors named | Which rivals appeared, for share of voice |
The “Sources listed” column is the one teams skip and later regret. It tells you who the model reads for your category, which is what you act on next. The pattern is consistent across studies: in our August 2026 measurement, Reddit was ChatGPT’s single most-cited source at about 30.3% of all citations, more than double Wikipedia at 13.8%, with TechRadar near 9.6% behind them. If Reddit, G2, and TechRadar keep showing up and your site does not, that is your off-page work, not a copy edit.
What does a real baseline look like?#
Most first baselines are humbling, and that is the honest starting point, not a failure. The first run shows the size of the gap, which prompts you lose, and who wins them instead. On a “best AI visibility tools” prompt, for example, ChatGPT led with Profound in position 1, then Semrush, Ahrefs, Peec AI, and Scrunch. For any brand missing from that list, that is not a morale hit; it is a target list. Every week after is measured against that first run, so it is the floor you climb from. The alternative, a vague sense that you “show up sometimes,” gives you nothing to act on.
When does a tool make more sense than a spreadsheet?#
A tool earns its place once you track many prompts across several engines and need share of voice and week-over-week deltas without a day of manual runs. Below that, a weekly spreadsheet is honest and cheap. The switch is about scale and consistency, not sophistication.
Common mistakes tracking AI mentions#
- Confusing mentions with citations. Record both. A recommendation with no link and a citation with no name are different signals.
- Rewording prompts between runs. Fixed wording is what makes the trend real. Freeze the set and grow it as a separate group.
- Switching models mid-track. A new model version moves your numbers on its own. Note the version and keep it fixed.
- Reading one session as truth. Answers vary by session; only the weekly trend is stable enough to act on.
- Tracking ChatGPT only. The engines rarely share sources, so a win on one is not a win on all. Track the set your buyers actually use.
- Skipping the sources column. Who the model reads is what you act on next. Log the cited domains every time.
How Prefer helps#
The manual method works, and for a small prompt set on one engine it is the right call. It breaks down at scale: dozens of prompts, five engines, a fresh session each time, and someone reading every cited source by hand, every week. That is the work Prefer automates. It runs your tracked prompts across ChatGPT, Gemini, Perplexity, Google AI Overviews and AI Mode (Claude on its Enterprise plan), records mentions and citations and positions, reads every source the answers cite, and computes your share of voice and how it changes over time.
We run this same loop on our own brand. To see where the AI engines stand on your brand today, run the free AI visibility checker or a full free AI visibility audit. If you also want to size the click side of the channel, read how to measure ChatGPT traffic, and for the numbers that tie visibility to revenue, how to measure AI search ROI.
People also ask
- What is the difference between a mention and a citation in AI answers?
- How do I know if ChatGPT mentions my brand?
- How do you measure share of voice in AI search?
- Can I track AI brand mentions in a spreadsheet?
- When should I use a tool to track AI mentions?
- How often should I check my AI brand mentions?