acceptodds
Under review as a conference paper at ICLR 2027

RETAIN AND REVEAL: MULTI-CONTINUATION TRAINING THROUGH AN OPTICAL MEMORY BOTTLENECK

Abstract

Large language model agents must preserve evidence for decisions whose requirements become clear only as new observations arrive. Long-context question answering provides a concrete setting: later passages can determine which earlier relations are needed. Under limited memory, these relations must survive both selective retention and compression into a readable representation. We introduce R&R, which trains one committed memory to support multiple source-grounded continuations. The writer commits without access to the continuation probes. Its memory is rendered once, and the reader answers the same question using the identical image paired with different continuations. Continuation rewards encourage retention of unresolved alternatives, while reader supervision trains evidence recovery through the optical channel. Paired memory skills coordinate evidence organization and interpretation, with failure-driven revisions validated on held-out histories under frozen model weights. Experiments on long-context QA benchmarks demonstrate competitive accuracy across settings under matched answer-time memory-input token allowances. Gains over textual memory approaches are more pronounced at smaller budgets, supporting the value of training bounded memory to preserve and recover evidence for subsequent reasoning. Controlled comparisons show mean-accuracy gains from continuation diversity, joint probe supervision, and paired skill revision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.