acceptodds
Under review as a conference paper at ICLR 2027

Remember Less, Adapt Longer: Bounded-State In-Context Reinforcement Learning with Recurrent Summary Memory

Abstract

In-context reinforcement learning (ICRL) enables agents to adapt to new tasks without parameter updates, relying on memory to retain what they learn from experience. For agents operating continually, existing approaches using memory trade off bounded cost against access to past experience. Approaches using growing memories, such as full-history attention and accumulating summaries, incur storage and per-decision costs that increase with task length. Fixed-size memories bound storage but restrict access: sliding windows discard older steps, while recurrent networks compress even recent interactions into a hidden state. To avoid this trade-off, we introduce Recurrent Summary Memory (RSM), a Transformer policy combining exact attention over a short-term working buffer with a fixed-capacity long-term summary of earlier experience. When the buffer fills, RSM rewrites the summary from its previous contents and buffered interactions, then clears the buffer, keeping deployment memory and per-decision cost in task length. On long-running partially observable benchmarks, frozen RSM policies sustain adaptation performance for eight times the training horizon using about 1% of full-history memory, and outperform all evaluated baselines at the final horizon, with a widening lead over growing memories. When the underlying task configuration changes mid-run, RSM adapts again and outperforms growing-memory and recurrent baselines. The policy relies on the summary: erasing it or blocking policy access removes its benefit, while swapping in another task's summary steers behaviour toward that task. Together, these results show that our fixed-capacity summary can retain task knowledge far beyond the training horizon.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.