AI visibility does not have a Search Console. There is no single dashboard every answer engine reports into, no universal ranking position, and no click-through rate you can pull from one API. Anyone measuring this well is building the measurement themselves, one buyer question at a time.
That is not a reason to skip it. It just means the measurement has to be built around evidence you can actually check: a real question, a real answer, a real citation or the lack of one. This guide covers how to build that measurement so it holds up, and what to leave out because it only looks like progress.
Start from real buyer questions, not one test prompt
A single prompt like "tell me about [brand]" tells you almost nothing, because it is not the question a real buyer types. Useful measurement starts with a library of the actual questions your buyers ask before choosing a provider: discovery questions ("best X for Y"), comparisons ("X vs Y"), reputation checks, specific feature questions, and transactional questions close to a decision.
Balance the library across those types. A library stacked with brand-named comparison questions will look artificially strong. A library of only broad discovery questions can miss the specific, high-intent questions where you actually need to show up.
Track cited, mentioned, or not found, separately for each engine
For every question, record three states per engine: cited (the engine names you and points to a source), mentioned (the engine names you with no clear citation), or not found (you do not appear at all). Collapsing this into one number hides the part that matters most, which is whether the mention comes with a source a buyer could actually click.
Run the same question across more than one engine. ChatGPT, Claude, Perplexity, and Gemini do not draw on the same sources or reasoning, and a brand can be cited clearly in one and absent in another for the same question.
- Cited. The engine names your brand and attributes the claim to a specific source, often with a link.
- Mentioned. The engine names your brand, but without a clear citation the buyer could follow.
- Not found. Your brand does not appear in the answer at all, even when it is a reasonable fit for the question.
Keep the evidence, not just a status
The status alone (cited, mentioned, not found) is only useful with the actual answer text and, where one exists, the source URL saved alongside it. That evidence is what lets you explain why a question failed, whether the model got a fact wrong, and whether a fix actually changed the wording of the next answer, not just a summary number that moved for unclear reasons.
Measure before and after a fix, on the exact same question
The clearest proof that anything you did mattered is a before-and-after pair on the identical question. Record the answer before you publish a change, publish the fix, wait long enough for that engine to reflect it, then run the exact same question again and compare the two answers side by side.
Give it real time. Different engines refresh on different schedules, and a single re-check the next day can just as easily catch a temporary answer as a durable one. Track a question across a few checks over a few weeks before treating a change as real.
What not to measure: vanity numbers that only look like progress
A few patterns show up often and mean less than they look like:
- A single composite score with no visible math. If you cannot see which questions and which engines produced the number, you cannot trust it or explain a drop.
- Content-completion percentages relabeled as visibility. Pages drafted, approved, or published is real progress, but it is a work-in-progress metric, not proof an engine is citing you.
- Raw crawl volume presented as AI citations. A crawler visiting your site is not the same as an engine citing your brand in an answer. Both are worth tracking, but they answer different questions.
- One good answer treated as a trend. A single cited answer on one engine, one day, is a data point, not a pattern. Watch it across engines and across repeat checks before calling it a result.
A measurement cadence you can actually keep up
- Build and maintain a real buyer-question library. Balanced across discovery, comparison, reputation, specific, and transactional questions, refreshed as your category and offering change.
- Probe multiple engines on a regular cycle. A weekly or steady rolling cadence beats one large one-time check, because engines update on their own schedules.
- Log the status and the evidence, every time. Cited, mentioned, or not found, plus the answer text and any source link, saved per question and per engine.
- Re-test the same question after any real change. New page, updated fact, or fixed structured data. The before-and-after pair is the actual proof of lift.
- Report the trend, not a single reading. One check is a snapshot. A few weeks of checks on the same library is a trend you can actually act on.
FAQ: measuring AI visibility
What does it mean for AI to "cite" a brand versus just "mention" it?
A citation means the engine attributes a specific claim to your brand and usually links to a source. A mention means your brand is named in the answer without a clear citation a buyer could follow. Both matter, but a citation is stronger evidence of trust in your source.
Is there one score I should track for AI visibility?
Be cautious of any single number with no visible methodology. Useful measurement tracks cited, mentioned, or not found per question and per engine, backed by the actual answer text, rather than one composite score that hides which questions and engines produced it.
How many buyer questions do I need to measure AI visibility well?
There is no universal number, but a handful of questions is not enough to see a real pattern. Aim for a balanced library across discovery, comparison, reputation, specific, and transactional questions so you are not only testing the easiest ones.
How often should I re-check AI visibility?
On a steady, repeated cadence rather than one large one-time check. AI engines update on their own schedules, so a rolling weekly or regular check across your question library shows real trends instead of a single moment in time.
Why does my brand show up in ChatGPT but not in Perplexity for the same question?
Different engines draw on different sources, retrieval methods, and reasoning. A brand that is cited clearly in one engine can be genuinely absent from another for the identical question, which is why measurement has to cover more than one engine to be trustworthy.