Surfer ran the same prompts through the major engines over and over and counted which citations came back. Perplexity held steadiest. ChatGPT churned the most. Google’s AI Overviews landed between the two (Surfer Academy). Same prompt, same week, different machines: one keeps time like a metronome, one pays out like a slot machine.
That spread changes what a single test run is worth. It also changes which fix you fund first.
How much do AI citations vary run to run?
Enough to make one run meaningless on the volatile engines. SparkToro had 600 volunteers push 12 identical prompts through three engines 2,961 times; the odds of two runs returning the same ordered brand list came out near 1 in 1,000. We covered the full math before. What Surfer’s data adds is that the churn has a shape. It differs by engine, and by citation source within an engine. A source cited in every ChatGPT run exists. So does one cited in every third run. Your tracker’s dashboard shows both as “cited.”
Is the variance just personalization?
No. Graphite sampled across users and found nothing worth the name; Ethan Smith: “I have not seen a significant difference in answers based on the person” (Ethan Smith x Niklas Buschner). The variance is the model’s own randomness, which means it forms a distribution you can measure. Graphite’s sampling rule, per Smith: “if you ask the question seven to ten times, you’ll get a general distribution of the citations that will appear for the next 100” (Surfer Academy). One run is one pull of the lever. Seven to ten runs is the payout table.
Why is a volatile citation worth half?
Because it misses a share of the answers your buyers see while costing the same effort to win. Surfer’s team, after sampling at volume, prices it without sentiment: a source that is “on and off, on and off” is “basically worth half” the effort of one that cites you every time (Surfer Academy). Half the presence. Full outreach bill.
Which points at the fastest payback in this whole corpus: flipping a volatile citation permanent. Surfer did it with one placement on an agency site drawing about 6,000 monthly Google visits, and its presence on the target prompt went from intermittent to “cited almost every time.” Their placement price sheet sharpens the point: sites cited about as often as a TechRadar-class domain sell placements at $375 and $700 against TechRadar’s near $20,000. The engines do not price citations. Publishers price prestige.
Where do you spend first?
On citations sitting near the flip threshold, measured on the engine you can trust at low sample counts. A source citing you in 5 of 7 runs is one explicit mention away from permanent. A source citing you in 1 of 7 needs the whole diet of third-party mentions, and that is a months-long program with a different budget line. Volatility data turns “get cited more” into a ranked queue.
Engine stability sets your measurement bill on top. A Perplexity reading stabilizes on few runs. A ChatGPT reading without repeated sampling is a screenshot with confidence it did not earn. Budget your runs where the machine is noisiest.
The ordering also hands you a cheap feedback loop. Test a change against the metronome first: a before-and-after reading on Perplexity settles fast, and the fix you validated there ships to every engine at no extra cost. Then sample ChatGPT at 7 runs to see whether the flip carried over into the noise. A stability score per citation source, tracked across audits, turns the slot machine into something you can at least count cards against.
Our spec follows this arithmetic. The free scan runs 5 buyer questions once per engine and labels the result a snapshot. Paid audits run each question 7 times per engine and link every finding to the captured answer. Mention rates come out as ranges. We never print a rank. Sample the distribution, find your 5-of-7 sources, and spend there first: run the free scan.