Nostimates
← All research
Methodology·2 min

The Citation Volatility Index: how much AI answers change between identical prompts

Across 400 prompts sampled 25 times each on six engines, only 18% of prompts returned an identical citation set on every repeat. Median citation-set overlap between two random samples of the same prompt was 62% on Google AI Overviews, 54% on AI Mode, 47% on ChatGPT and 81% on Perplexity — which means any single reading of an AI answer misstates a brand's position roughly half the time.

Method
400 prompts × 25 samples × 6 engines, US market, English, logged-out collection
Sample
60,000 answers
Citation volatility across AI answer engines

Why we ran this

Every AI visibility tool on the market shows a number. Almost none of them show how much that number would have moved if the prompt had been asked again ten minutes later. We wanted an empirical floor under a claim we make constantly: one reading is not a measurement.

Method

We selected 400 commercial-informational prompts across eight verticals (SaaS, finance, travel, health, ecommerce, legal, education, B2B services). Each prompt was run 25 times per engine over a 72-hour window, logged-out, from residential egress in the United States with hl=en. We recorded the ordered citation set and the set of brands named in the answer body.

Engines: Google AI Overviews, Google AI Mode, ChatGPT, Perplexity, Microsoft Copilot, Baidu ERNIE (CN prompts excluded from the headline figures).

What we found

Engine Identical set on all 25 samples Median pairwise overlap Median distinct sources across 25 samples
Perplexity 34% 81% 9
Google AI Overviews 21% 62% 14
Microsoft Copilot 19% 60% 13
Google AI Mode 11% 54% 21
ChatGPT 8% 47% 26

Three things stand out.

Volatility is engine-specific, not universal. Perplexity's structural citation model produces answers that are close to reproducible. ChatGPT's are not: the median prompt surfaced 26 distinct sources across 25 samples, meaning most sources appeared once.

Brand mentions are more stable than citations. A brand named in the answer text appeared in a median of 76% of samples where it appeared at all, versus 58% for a specific URL. Brand-level measurement is meaningfully less noisy than URL-level measurement.

Volatility rises with query breadth. Prompts with fewer than six words showed 14 points more spread than prompts over twelve words. Head terms are the least reproducible thing you can track.

The practical consequence

At n=1, the 95% confidence interval on a citation share estimate is effectively the whole range. Around a 30% share, the Wilson interval is roughly ±25 points at n=10, ±17 at n=25, ±12 at n=50 and ±9 at n=100. Choose your sample size against the size of the change you need to detect, not against your budget.

What this means for reporting

If your report says "we are cited in AI Overviews", the honest version of that sentence includes a sample size and an interval. If a client's competitor "appeared" this month, check whether the appearance survives resampling before you write it into a deck.

We publish the per-engine volatility figures we use for our own interval calculations, and they are baked into every run our API returns.

Evidence

Raw answers, per-cell sample sizes, collection dates and parser versions behind every figure here are available to customers on request. We publish the method so the numbers can be checked rather than taken on faith.

See this for your brand

Every figure here comes from the same measurement behind our dashboard and API. Tell us what you want measured and we'll show you where your brand stands.

No newsletter. One reply from a human.