A 40-day measurement of 55,393 trending queries finds Google AI Overviews activate on 13.7% of searches (64.7% for question-form queries) and that 11.0% of 98,020 atomic claims are unsupported by their cited sources
Synopsis
Using a Puppeteer crawler on AWS Lambda, the authors ran a 40-day longitudinal measurement (March 13–April 21, 2026) over 55,393 trending Google queries across 19 topical categories and found that AI Overviews activate on 13.7% of queries overall (64.7% for question-form queries), that cited domains are more credible than co-displayed first-page results yet 29.8% do not appear on the first page at all, that 11.0% of 98,020 atomic claims are unsupported by the cited pages with omission as the dominant failure mode, and that at least 50.6% of cited pages carry display advertising.
Figure 1: Our approach to characterizing the Google AIO ecosystem: (1) We begin by extracting top search queries from Google Trends. (2) We develop a bespoke crawler to capture search results for these queries, including their AIOs and associated metadata (e.g., reference links). We then crawl the referenced pages to extract their content and detect the presence of advertisements. (3) We evaluate the quality of AIOs by assessing the credibility of cited references and the consistency of AIO claims with source content. Finally, we analyze the potential economic impact of AIOs on the web ecosystem.
· Page 5Interpretation
AI Overview activation is highly structured: 13.7% overall, 64.7% for question-form queries versus 9.5% for non-question queries (a 6.8 times difference), and a 13-fold spread across categories from 3.5% in Beauty & Fashion to 46.1% in Hobbies & Leisure, with Politics (7.5%) and Law & Government (9.6%) below average. Prior work was either narrow in domain or relied on fixed query sets collected before AIOs reached their current scale; this study characterizes triggering across a diverse, naturalistic trending-query sample and isolates the independent effects of query phrasing and length (non-question queries rise from 9.9% at one word to 38.7% at six or more). Based on 55,393 queries and 7,583 AIO activations over 40 days; the question versus non-question gap is tested with a chi-square statistic (χ² = 10,002.2, p < 10⁻³⁰⁰), and category variation comes from stratified counts across 19 categories.
AIO-cited domains are systematically more credible than co-displayed first-page results (mean PC1 0.732 versus 0.645, a gap of +0.087, 95% CI [0.085, 0.089]), yet 29.8% of cited domains do not appear anywhere on the corresponding first page, indicating a source pool distinct from Google's own ranking algorithm. This runs counter to prior work using categorical credibility taxonomies that found AIOs drawing on lower-quality sources; the authors use a continuous PC1 domain-quality score and additionally quantify overlap between the two source pools, showing AIO selection is not a simple re-ranking of the first page. 37,020 AIO references matched to PC1 versus 159,752 first-page URLs (Welch's t = 80.9, p ≪ 0.001); 14 of 19 categories remain significant after Bonferroni correction with the direction favoring AIO references; off-page references are higher quality (PC1 0.758 versus 0.724; UGC share 3.4% versus 18.5%).
Decomposing 7,491 verifiable AIOs into 98,020 atomic claims, 88.97% are consistent with their cited sources and 11.03% are not: Omitted 6.98%, Incorrect 2.66%, Ambiguous 1.39%; omission rather than fabrication is the dominant failure mode, and source quality and claim fidelity are largely independent (r ≈ 0.045). Google states AIOs 'generally don't hallucinate in the ways that other LLM experiences might'; this study verifies each claim against the full text of every cited reference using a five-label scheme, producing a quantified inconsistency rate and showing that improving sourcing alone is unlikely to reduce unsupported claims. The pipeline was validated against human annotation: claim extraction reached precision 98.26%, recall 83.23%, F1 90.12%; the verifier matched human labels on 98 of 100 stratified verdicts with 95.6% overall accuracy weighted by label frequency, and never wrongly flagged a grounded claim as Incorrect or Omitted.
Of 61,212 cited URLs, 30,994 (50.63%) display visible ads, while 164 of 7,583 AIO-bearing SERPs (2.16%) also carry a Google sponsored search ad and 39 (0.51%) place a sponsored slot above the AIO block, with the two occupying disjoint page regions. Prior work documented click and traffic declines from AIOs but did not characterize the monetization structure of the cited pages themselves, nor whether Google's own ad inventory is preserved on the same page; this study measures both sides together. Based on DOM parsing of all 61,212 cited URLs (EasyList rules) and sponsored-ad detection across all 7,583 AIO-bearing SERPs over the 40-day window; the authors note 50.63% is a lower bound because the 14.2% of references pointing to social and video platforms were not crawled and are treated as ad-free.
Perspective
The work provides a reusable baseline for continued measurement: activation, source quality, claim fidelity, and ad exposure can all be re-run on the same crawler and verification pipeline, making it useful to regulators, independent auditors, and news and platform researchers tracking how the system evolves. The results apply to US-localized Google Search, entered through trending queries, within the March 13–April 21, 2026 window; the authors also note that commercial-intent queries are underweighted in their sample, which is precisely where sponsored ads are most prevalent. For publishers, the directly usable output is the ad share and category distribution of cited pages; for readers, it is the structural fact that question-form and longer queries are far more likely to trigger an AIO.
The authors list several open questions: the study measures what AIOs say and cite, not what users do with them—whether errors propagate into belief, whether corrections are sought, and whether effects are more severe for high-stakes topics like health and political information remain unanswered; how AIO behavior varies across user profiles, personalization states, and geographic contexts is also unknown. Methodologically, social and video platforms were not crawled, so the authors treat the 11.0% inconsistency rate as an upper bound and estimate a residual inconsistency rate of roughly 5.3% under the most generous assumption. The low fidelity in Climate and Jobs & Education is attributed to real-time values (weather, school closures) shifting between synthesis and crawl, a pipeline-level failure mode rather than a genuine AIO quality problem, so the combined 9.64% Omitted and Incorrect rate should be read as a conservative ceiling on substantive unfaithfulness. Finally, the 2.16% sponsored-ad co-occurrence rate reflects a sample dominated by informational and trending queries; the commercial-intent query space still needs separate measurement.
