RETAIN AND REVEAL: MULTI-CONTINUATION TRAINING THROUGH AN OPTICAL MEMORY BOTTLENECK
Abstract
Large language model agents must preserve evidence for decisions whose requirements become clear only as new observations arrive. Long-context question answering provides a concrete setting: later passages can determine which earlier relations are needed. Under limited memory, these relations must survive both selective retention and compression into a readable representation. We introduce R&R, which trains one committed memory to support multiple source-grounded continuations. The writer commits without access to the continuation probes. Its memory is rendered once, and the reader answers the same question using the identical image paired with different continuations. Continuation rewards encourage retention of unresolved alternatives, while reader supervision trains evidence recovery through the optical channel. Paired memory skills coordinate evidence organization and interpretation, with failure-driven revisions validated on held-out histories under frozen model weights. Experiments on long-context QA benchmarks demonstrate competitive accuracy across settings under matched answer-time memory-input token allowances. Gains over textual memory approaches are more pronounced at smaller budgets, supporting the value of training bounded memory to preserve and recover evidence for subsequent reasoning. Controlled comparisons show mean-accuracy gains from continuation diversity, joint probe supervision, and paired skill revision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.