verisible AI visibility, measured honestly
← Field notes
2026-07-05 · geo · measurement · ai-search

There is no AI rank. Here's the math

600 volunteers ran 12 identical prompts through ChatGPT, Claude, and Google’s AI a total of 2,961 times. The odds of two runs returning the same brand list in the same order came out to roughly 1 in 1,000 (SparkToro). A single “AI rank” is a coin flip you watched land once.

There is no stable AI rank because AI answers are non-deterministic. The same prompt returns different brands, in a different order, at a different list length, run to run. Rand Fishkin and Patrick O’Donnell measured under a 1-in-100 chance of an identical list and about 1-in-1,000 for identical order. Any tool selling you a fixed rank position is reporting one sample and calling it a law.

This is not a rounding error a vendor can smooth over. It is the core behavior of the system. Temperature, personalization, and session memory all inject variance, so “you rank #3 in ChatGPT” describes a moment that has already passed.

What did the study actually measure?

Twelve prompts asking for brand recommendations across ordinary categories: chef’s knives, headphones, cancer care hospitals, digital marketing consultants, science fiction novels. Nearly 3,000 runs across three engines. The team then checked how often the answers repeated (Search Engine Land).

Three things varied almost every time: which brands appeared, the order they appeared in, and how many the list contained. Claude held together marginally better on ordering than ChatGPT or Google, and marginally is doing real work in that sentence. The headline stands across all three: identical ordered lists showed up in about 1 run in 1,000 (Search Engine Journal).

The uncomfortable part for the tracking industry: those dashboards showing your brand at “position 4 in ChatGPT this week” are sampling a distribution once and printing the result as if it were a fact. Do it again tomorrow and position 4 might be position 1, position 8, or absent. The tool will show you a new number with equal confidence and no error bar.

Why does the number move so much?

Three forces, stacked. Temperature is deliberate randomness in how the model picks its next word, so even a fixed context produces varied output by design. Personalization means the answer bends to whoever asked, their history, their phrasing, their account. Session memory means the same account can get different answers depending on what it discussed earlier.

Add query fan-out on top. The engine decomposes your question into sub-queries generated fresh each time, so the retrieval underneath the answer shifts run to run before the model even starts writing. The wonder is not that answers vary. The wonder is that anyone expected a slot machine to return a stable leaderboard.

We see this on our own scans constantly. Running the same five-question check against a single site on two different days, we got 2 of 5 mentions the first time and 3 of 5 the second. The site did not change. Our questions did not change. A tool reporting the first run would have told the client “you appear in 40% of answers.” The second run says 60%. Both are true snapshots of an unstable process, and neither is a rank.

So what do you measure instead?

Visibility percentage across a large or repeated sample. The same SparkToro analysis that shredded single-run ranking found that visibility, the share of answers mentioning your brand across dozens to hundreds of runs, holds up as a reasonable metric. The variance that destroys a single rank averages out into a stable rate when you sample enough.

That is the whole methodology, and it is not complicated: fix a question set, run it many times across the engines you care about, and report the range of how often you appear, not a position. A brand that shows up in 55 to 65% of answers over 100 runs has a real, defensible visibility figure. A brand reported at “rank #3” has a screenshot.

We built our checker to report ranges over multiple runs and to label single-run results as snapshots, never as ranks, because the alternative is lying to a client with a straight face. The disciplined version of this, weighted and tracked over time, is share of voice in AI answers, which is the only AI visibility metric that survives the variance this study documents.

The practical next step costs nothing and takes ten minutes. Pick five questions your buyers ask, run them through a live engine twice on two different days, and watch the answer change under your own eyes. Our free scan runs the first pass and shows the citations verbatim: see it move. Once you have watched the number wobble yourself, no vendor’s fixed-rank dashboard will ever look the same.

Sources

See it on your own site. Five buyer questions to an AI, and we show you whether you're in the answers.
Start tracking free