acceptodds
Under review as a conference paper at ICLR 2027

Beyond K-step Histories: Learning to Remember Reward Relevant Observations for Non-Markovian RL

Abstract

Reinforcement learning with partial observations can require retaining information across long temporal delays, even in controlled environments. We study how computationally limited agents can manage this information when only a small subset of past observations is relevant for decision-making. The predominant approach, sliding windows (also called frame stacking), maintains a window (stack) of the most recent observations, but the required stack size grows with the degree of non-Markovian dependencies, leading to prohibitive computational and memory requirements for both inference and learning. We build on the observation that many environments that are highly non-Markovian over time depend causally on only a small subset of past observations. This motivates meta-algorithms that maintain small adaptive memory stacks, enabling the representation of long-term dependencies while processing fewer observations per step. We propose Adaptive Stacking, a meta-algorithm with convergence guarantees, and quantify its computational and memory savings for MLP, LSTM, and Transformer-based agents. Using controlled memory tasks with varying degrees of non-Markov dependencies, we show that Adaptive Stacking learns to discard memories that are not predictive of future rewards while retaining important experiences.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.