AI Overviews and AI Mode disagree on 6 of every 10 queries
On 1,000 identical queries run against both Google AI Overviews and Google AI Mode, the two surfaces shared a majority of their citations on only 39% of queries. Median overlap was 41%, and on 12% of queries the two surfaces shared no citations at all — so treating them as one 'Google AI' number hides more than it reveals.
- Method
- 1,000 queries run against AI Overviews and AI Mode within a 15-minute window, US, English
- Sample
- 20,000 answers

The question
Google runs two AI answer surfaces on the same index. If they largely agree, tracking one is enough. If they don't, any tool that reports a single "Google AI visibility" figure is averaging two different things.
Method
1,000 queries, balanced across informational and commercial intent, run against both surfaces within a fifteen-minute window to control for index drift, ten samples each, US residential egress, English. Overlap is measured as Jaccard similarity on the deduplicated domain set.
Results
| Metric | Value |
|---|---|
| Median citation overlap (domains) | 41% |
| Queries where the surfaces shared a majority of citations | 39% |
| Queries with zero shared citations | 12% |
| Median distinct domains, AI Overviews | 5 |
| Median distinct domains, AI Mode | 11 |
AI Mode's query fan-out is the mechanism. Decomposing one question into up to sixteen parallel searches pulls in sources that never rank for the literal query, so the citation set is wider and less predictable from classic ranking data.
Divergence is worst on commercial queries. Product comparison and "best X for Y" prompts had a median overlap of 29%, versus 52% for definitional queries. The queries agencies care about most are the ones where the two surfaces agree least.
Winning one does not predict winning the other. Of brands cited in AI Overviews for a query, only 44% were also cited in AI Mode for the same query. The correlation is real but weak enough that substituting one for the other is a reporting error.
Reporting implication
What to do about it
Track both, weight by the audience you actually care about, and re-baseline when the merge reaches your market rather than assuming continuity in the trend line.
Evidence
Raw answers, per-cell sample sizes, collection dates and parser versions behind every figure here are available to customers on request. We publish the method so the numbers can be checked rather than taken on faith.
See this for your brand
Every figure here comes from the same measurement behind our dashboard and API. Tell us what you want measured and we'll show you where your brand stands.