verisible AI visibility, measured honestly
← Field notes
2026-07-17 · geo · measurement · ai-search

The people selling AI rankings keep saying they don't exist

Tim Soulo is the CMO of Ahrefs, a company that sells an AI visibility product called Brand Radar. Here he is on his own podcast: “Anyone who is claiming that they have accurate volume data for prompts, they’re just essentially lying to you” (Ahrefs Podcast). He sells the category. On air, he filed a confession against it.

He had company in the room. Ethan Smith runs Graphite, the agency behind the canonical AEO playbook, and he itemized what a custom-prompt tracker does not know: the prompts people use, their volume, whether your category is the primary subject of the prompt, whether the query is search or generative. His verdict on the stacked unknowns: “all these errors just having this wildly, extremely wide confidence interval that’s almost meaningless.”

The confessions extend across the vendor corpus. Open Forge, which sells visibility tracking, on its own product class: “anyone telling you it’s exact is not being entirely truthful.” iPullRank, which sells GEO consulting, tells clients to “expect precision, not accuracy.” Graphite counted 60 answer-tracking tools on the market, judged them “all roughly the same,” and advised buyers to pick the cheapest. Smith adds that a tracker is something you could “write in a day.” Four sellers, one message, delivered while the invoices kept going out.

Why would vendors undercut their own dashboards?

Because the number they sell does not hold still, and they know it from their own data. SparkToro had 600 volunteers run 12 identical prompts through three engines a total of 2,961 times, across categories as plain as chef’s knives and headphones. The odds of two runs producing the same ordered brand list came out under 1 in 1,000 (SparkToro). A rank needs a repeatable query underneath it, and Conductor, a search analytics company, priced that assumption: “What’s the monthly search volume of a prompt? Most likely, the number is one.”

The prompt sets are the second hole. Trackers sell packages of 25 or 50 hand-picked prompts, and Soulo asked the obvious question in public: “How can you rely on 25 manually selected custom prompts to gauge your AI visibility?” Then he turned the audit on his own product: “Full transparency. I’m not happy with how we’re using search volumes to extrapolate into the prompts that we use in Brand Radar.” Smith, after talking to multiple AEO companies about the panel data that would fix this: “Panel data today is not ready.” The sales page sells positions. The podcast admits distributions. You are buying the sales page and receiving the podcast.

What does honest AI visibility measurement look like?

Disclosed samples, ranges instead of points, and evidence you can open. The vendors wrote this spec themselves, in their own defense. Smith’s sampling prescription: ask the same question many times, “multiple runs, like 10 runs of the same question on each surface,” and read the output as a probability distribution. Graphite’s own research puts stable sampling at 7 to 10 runs per prompt. The posture worth copying comes from Graphite’s traffic study, shipped with its raw data: “everything is reproducible. So you don’t need to trust us.”

That is the standard we hold ourselves to, since the alternative is joining the confession file. Our free scan asks 5 questions your buyers would ask, runs them once, and labels the result a snapshot, with the engine’s citations shown in full. The paid audit runs every question 7 times per engine and reports mention rates as ranges. Every finding links to the captured answer it came from. No rank appears anywhere, for reasons the math already covers.

When every seller in a market volunteers the same weakness in the product, that is the one claim you can take at face value. So put two questions to any visibility vendor before you pay: how many runs per prompt, and can I see the raw answers. A vendor with good methodology answers both in one breath. A vendor selling a rank changes the subject, and the people quoted above have already told you what the subject change means. See what a disclosed-sample scan looks like.

Sources

See it on your own site. Five buyer questions to an AI, and we show you whether you're in the answers.
Start tracking free