verisible AI visibility, measured honestly
← Field notes
2026-07-05 · geo · measurement · share-of-voice

Share of voice in AI answers: the only metric that survives scrutiny

Ahrefs’ Brand Radar tracks brand mentions across roughly 391 million monthly prompts spanning six engines: AI Overviews, AI Mode, ChatGPT, Perplexity, Gemini, and Copilot (Ahrefs Brand Radar). The scale is the point. You cannot read AI visibility off one answer, so the only figure worth reporting is a rate measured across many.

AI share of voice is the percentage of AI answers, across your target questions and chosen engines, that mention or recommend your brand versus competitors (Digital Applied). It survives scrutiny because it is built from a sample, not a single run, so the run-to-run randomness that makes “AI rank” meaningless averages into a stable number you can defend to a board.

Everything else marketers try to report about AI visibility falls apart under a second run. Share of voice, done with enough samples and disclosed methodology, does not.

Why does sampling beat a rank number?

Because AI answers are non-deterministic and a single run captures noise. The SparkToro study of 2,961 prompt runs found under a 1-in-1,000 chance of two answers repeating the same ordered brand list. A rank from one answer is a screenshot of a slot machine.

Share of voice sidesteps this by treating each answer as one draw from a distribution. Ask your question set 100 times and count how often you appear. The individual answers scatter wildly; the aggregate rate is stable. That is the entire statistical argument, and it is why every serious measurement voice in this space converges on sampled share rather than position. We took apart the single-number fiction in detail in there is no AI rank.

The catch is that sampling only works if you disclose it. A share-of-voice figure without a stated run count is as untrustworthy as a rank. “You appear in 60% of answers” means nothing until you add “across 100 runs of 12 questions on ChatGPT and Gemini, sampled over five days.” The methodology is the metric. Hide it and you are back to dashboard theater.

What does an honest share-of-voice metric contain?

Three ingredients, all declared. The query set: the branded, competitor, and problem-solution questions your buyers actually ask. The engines: which answer surfaces you monitor, because Copilot and Perplexity and ChatGPT do not agree with each other. The scoring rule: whether a brand was merely mentioned, cited as a source, or explicitly recommended, weighted so a first recommendation counts more than a buried aside (Alex Birkett).

That weighting matters more than it looks. Being named as the top recommendation in an answer is worth more than appearing in a list of eight. A flat mention count treats those the same and flatters brands that get name-dropped without being endorsed. A good share-of-voice score separates presence from prominence.

And it reports a range, not a point. Because the underlying answers vary, the honest output is “42 to 51% across the sample,” which tells a reader both your level and your volatility. A brand pinned at a single percentage to two decimal places is a brand whose tool is pretending the variance does not exist.

Where does citation dominance fit in?

Mention rate and citation dominance are different signals, and watching both is where this gets useful.

We see the split in our own scans. Running one site through two independently generated question sets, the surrounding brand mentions shuffled between runs, wobbling the way non-determinism guarantees. But one source, fertilityclinicsabroad.com, held: cited in 5 of 5 answers in the first set and 4 of 5 in the second. Its citation presence barely moved while the mention noise churned around it.

That stability is the real prize. Mention rate tells you how often the model says your name. Citation dominance tells you how often the model leans on you as a source, and a source cited in nearly every answer owns the category in a way no mention count captures. When we watched our own five-question check swing from 2 of 5 mentions to 3 of 5 on different days, the mention rate moved and the dominant citations did not. Tracking both is how you tell a genuine position from a lucky run.

The practical build is not exotic. Fix a question set your buyers would actually type. Pick two or three engines. Run each question enough times that the rate stabilizes, weight recommendations above mentions, and report the range with the run count attached. That is a share-of-voice metric a CMO can take into a board meeting without flinching, because every number in it has a stated sample behind it.

Start with the smallest honest version. Five real questions, one live engine, the actual citations read back to you. Our free scan does exactly that: run your first sample. Then decide whether the number is worth measuring at scale.

Sources

See it on your own site. Five buyer questions to an AI, and we show you whether you're in the answers.
Start tracking free