acceptodds
Under review as a conference paper at ICLR 2027

Train Short, Act Long: Multistability and the Horizon Independence of Recurrent Policies

Abstract

In partially observable reinforcement learning (RL), a decision often depends on a cue seen many steps earlier, and training at such horizons is costly. Ideally, a policy trained on short horizons would remain optimal at arbitrarily long ones, making the training horizon a matter of convenience. We formalize this property as *delay-horizon independence* (DHI): optimality at every horizon, at a memory, computation and precision cost that is bounded in the horizon. We study agents that carry the cue in a recurrent memory of fixed size, whose cost does not grow with the delay. This paper characterizes DHI through the dynamics of that memory, in theory and on benchmarks. Although quantified over all horizons, DHI reduces to a property of a single map: a recurrent policy is DHI if and only if it admits a continuous readout of the memory, invariant under the idle dynamics and separating the cues. Monostable memory then precludes DHI at any parameter values, whereas multistable memory enables it. Monostability covers all contracting affine recurrences, hence state space models and gated linear recurrent neural networks (RNNs). Under arbitrarily small perturbations of the idle dynamics, multistability is the only mechanism supporting DHI at bounded memory. On the T-maze, the effective horizon of monostable models resists hyperparameter search: selected horizons fall short by up to when retrained. Multistable models trained on short horizons stay optimal up to , on the T-maze and on MiniGrid. Current parallelizable RNNs thus trade horizon independence for scalability: parallelizable multistable architectures are the direction this points to.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.