The Citation Volatility Index: how much AI answers change between identical prompts
Across 400 prompts sampled 25 times each on six engines, only 18% of prompts returned an identical citation set on every repeat. Median citation-set overlap between two random samples of the same prompt was 62% on Google AI Overviews, 54% on AI Mode, 47% on ChatGPT and 81% on Perplexity — which means any single reading of an AI answer misstates a brand's position roughly half the time.
- Method
- 400 prompts × 25 samples × 6 engines, US market, English, logged-out collection
- Sample
- 60,000 answers

Why we ran this
Every AI visibility tool on the market shows a number. Almost none of them show how much that number would have moved if the prompt had been asked again ten minutes later. We wanted an empirical floor under a claim we make constantly: one reading is not a measurement.
Method
We selected 400 commercial-informational prompts across eight verticals (SaaS, finance, travel, health, ecommerce, legal, education, B2B services). Each prompt was run 25 times per engine over a 72-hour window, logged-out, from residential egress in the United States with hl=en. We recorded the ordered citation set and the set of brands named in the answer body.
Engines: Google AI Overviews, Google AI Mode, ChatGPT, Perplexity, Microsoft Copilot, Baidu ERNIE (CN prompts excluded from the headline figures).
What we found
| Engine | Identical set on all 25 samples | Median pairwise overlap | Median distinct sources across 25 samples |
|---|---|---|---|
| Perplexity | 34% | 81% | 9 |
| Google AI Overviews | 21% | 62% | 14 |
| Microsoft Copilot | 19% | 60% | 13 |
| Google AI Mode | 11% | 54% | 21 |
| ChatGPT | 8% | 47% | 26 |
Three things stand out.
Volatility is engine-specific, not universal. Perplexity's structural citation model produces answers that are close to reproducible. ChatGPT's are not: the median prompt surfaced 26 distinct sources across 25 samples, meaning most sources appeared once.
Brand mentions are more stable than citations. A brand named in the answer text appeared in a median of 76% of samples where it appeared at all, versus 58% for a specific URL. Brand-level measurement is meaningfully less noisy than URL-level measurement.
Volatility rises with query breadth. Prompts with fewer than six words showed 14 points more spread than prompts over twelve words. Head terms are the least reproducible thing you can track.
The practical consequence
What this means for reporting
If your report says "we are cited in AI Overviews", the honest version of that sentence includes a sample size and an interval. If a client's competitor "appeared" this month, check whether the appearance survives resampling before you write it into a deck.
We publish the per-engine volatility figures we use for our own interval calculations, and they are baked into every run our API returns.
Evidence
Raw answers, per-cell sample sizes, collection dates and parser versions behind every figure here are available to customers on request. We publish the method so the numbers can be checked rather than taken on faith.
See this for your brand
Every figure here comes from the same measurement behind our dashboard and API. Tell us what you want measured and we'll show you where your brand stands.