You measure AI visibility by triangulating three things, never one: repeated prompt sampling reported as a range, deterministic analytics from GA4 and your CRM, and direct surveys asking customers whether AI sent them. Any single layer lies by omission. Together they give you a defensible picture. A “you rank #3 in AI” number is not measurement, it is theater.
Start with the fact that breaks most tools. AI answers are non-deterministic. Ask the same question twice and you get different brand lists. The Fishkin and O’Donnell study ran 12 prompts across 2,961 runs on ChatGPT, Claude, and Google’s AI, and found less than a 1-in-100 chance that two runs of the same prompt returned the same ordered list. Fishkin’s own summary: any tool that hands you a “ranking position in AI” is, in his word, baloney.
Why is a single AI check meaningless?
Because you are sampling a distribution, not reading a fact. We ran the same 5-question visibility check against the same site on two different days. First run: 2 out of 5 questions mentioned the brand. Second run: 3 out of 5. Nothing about the site changed between them. The score moved because the model rolled its dice again. If we had reported either number as “your AI visibility,” we would have been reporting the weather and calling it the climate.
The fix is not a better single check. It is more checks. Run your prompt set at least three times, ideally more, and report the result as a range: “mentioned in 2 to 3 of 5 answers across three runs on two days.” A range is honest about the variance. A point estimate pretends the variance is not there, which is the exact thing SparkToro’s researchers warn against. This is also why there is no such thing as an AI rank, and why anyone selling you one is selling a fiction with a decimal point.
What does layer one actually cost?
Less than you would guess, which is why doing it properly has no excuse. We keep a cost ledger on every check. A 5-question grounded scan on Gemini’s free tier costs exactly $0.00. The same five questions run through OpenAI’s native web search cost about $0.032 per question, receipted in the ledger. So a rigorous multi-run sample, the thing that turns a lie into a measurement, is a rounding error. The reason most tools report a single flimsy number is not cost. It is that a single number demos better than a range.
What do the other two layers add?
Sampling tells you whether AI systems mention you. It cannot tell you whether that mention did anything. For that you need the deterministic layer: GA4 and your CRM, watching for sessions and pipeline that trace to AI sources. This layer is where the money shows up, and it is also badly broken, which is the subject of its own accounting. AI referral traffic routinely lands in GA4 as “direct” or unattributed, so the deterministic layer undercounts by default. You correct for it, you do not trust it raw.
The third layer is the one everyone skips because it is unglamorous: ask people. Add “how did you hear about us” to your signup or checkout, with an explicit “an AI assistant (ChatGPT, Gemini, Claude, Perplexity)” option. Survey data is soft and self-reported, and it is also the only layer that captures the customer who asked ChatGPT about your category, decided on you, and then typed your name straight into the address bar leaving no referral trail at all. That path is invisible to layers one and two. It only shows up if you ask.
Each layer has a characteristic failure. Sampling measures mentions that may convert nothing. Analytics measures conversions but mislabels their origin. Surveys measure perceived origin but rely on fuzzy human memory. Stacked, the failures do not overlap, so where the three agree you can actually believe the number.
What does honest reporting look like?
It discloses its own method. Say how many questions you asked, how many times you ran them, and on which engine, every time. State ranges, not points. Label a single run as a snapshot, not a rank. When the analytics undercount AI traffic, say so and show the survey number beside it. None of this makes the report weaker. It makes it the one report in the buyer’s inbox that is not pretending the measurement is cleaner than it is.
Pick five questions a real customer would ask an assistant about your category, run them three times over two days on a grounded engine, and write down the range. That is your baseline, and it costs close to nothing. Our free scan runs exactly this kind of multi-question check and reports what it found as evidence you can read, not a rank we invented.