Skip to content

PAPER

Google AI Overviews at scale — when it appears, what it cites, and whether the citation holds

On which queries do Google AI Overviews appear, which sources do they cite, and do those citations actually support the claims?

Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact

Authors
Haofei Xu · Umar Iqbal · Jacob M. Montgomery
Affiliation
Washington University in St. Louis
Venue
arXiv preprint
Submitted
2026-05-13
arXiv
arXiv:2605.14021
We verified
2026-08-24

WHAT THE PAPER SAYS

They appeared especially often on question-shaped queries; the sources cited were more trustworthy on average than ordinary search results; but a meaningful share of the generated claims was not supported by the pages cited. Across 55,393 queries observed over 40 days, the activation rate was 13.7% overall and 64.7% on question-shaped queries. Of 98,020 atomic claims, 11.0% lacked support from the cited page, and the authors state that crawl exclusions may have inflated this, putting the lower bound at about 5.3%.

Read the original on arXiv ↗

METHOD

How it was measured

Read the conditions before the numbers. The same figure means something different under a different sample or environment.

Queries
55,393 trending queries across 19 topic categories
Period
40 days (2026-03-13 to 04-21)
Claim verification
AI Overviews were split into self-contained factual units (pronouns replaced with entity names, duplicates removed) and classified as Clear / Vague / Ambiguous / Incorrect / Omitted
Verification accuracy
95.6% pipeline accuracy against human annotation

FINDINGS

What came out

Activation rate
13.7% of all queries, 64.7% of question-shaped queries. Markedly lower on politically sensitive topics
Claim fidelity
11.0% of 98,020 atomic claims were unsupported by the cited page
⚠️ The condition on that 11.0%
The authors excluded Facebook, Instagram, Reddit, TikTok, X and YouTube from crawling, which may have inflated the figure, and so proposed a lower bound of about 5.3%. The real value should be read as a 5.3–11.0% range
Source quality
Mean trust score of domains cited by AI Overviews 0.732 vs 0.645 for first-page results on the same screen — +0.087 on a 0–1 scale. Cited sources skew more trustworthy than ordinary results
Publishers
Well over half of the pages cited by AI Overviews carried display advertising

LIMITATIONS

Limitations the authors state themselves

Not our criticism — this is what the authors wrote in the paper.

  • Social media (Facebook, Instagram, Reddit, TikTok, X, YouTube) was excluded from crawling, which may have inflated the unsupported-claim rate — the lower bound is about 5.3%.
  • For 262 paywalled pages (<1%), only the visible portion was collected.
  • Queries that change in real time, such as weather or school closures, were captured hours or days later, by which point values had already shifted.
  • The low share of commercial-intent queries limits conclusions about ad co-presence.
  • There is no dedicated Limitations section; limitations are handled throughout the text.

NAVIRANG'S READING — NOT THE PAPER'S CONCLUSION

Two numbers here will come up constantly in practice. One is **64.7% activation on question-shaped queries** — a numeric reason to build documents that answer questions rather than company profiles and product lists. The other is that **5.3–11.0% of claims are not supported by the cited source**. That puts a scale on something we keep meeting in correction work: a source link being attached does not mean the sentence came from that source. Quoting only the 11.0% would misrepresent the authors. And all of it is US Google — nothing here has been verified for Naver AI Search or Korean-language queries.

If you want your own brand's numbers rather than a paper's

Every figure here came from someone else's sample. Send us a URL and we ask all 7 answer engines directly and measure yours.

We reply within one business day.

Free audit Call Email Blog