The Low-Rank Hankel Hypothesis in Reinforcement Learning
Abstract
Identifying and exploiting structure in reinforcement learning (RL) can guide exploration and improve sample efficiency, particularly when environment interactions are costly. We investigate the *low-rank Hankel hypothesis*: value-function sequences evaluated along trajectories of high-performing policies exhibit low effective Hankel rank, and can therefore be approximated by low-order linear recurrences. We provide empirical evidence for this structure in learned critic outputs across Classical Control, MuJoCo, and Atari. Theoretically, we show that such structure can arise under policies inducing geometrically mixing state chains. Motivated by these findings, we introduce Q-Hankeling, a simple method that regularizes Q-function learning using short trajectory chunks sampled from replay. The method learns linear recurrence coefficients jointly with the critic and penalizes recurrence residuals within these chunks. This provides a temporal-predictability inductive bias for value estimation, without explicitly constructing Hankel matrices or computing singular-value decompositions. We demonstrate how Q-Hankeling can be integrated into value-based methods, such as DQN and Efficient Rainbow, and actor–critic methods, such as TD3 and SAC. On Atari-100k, Q-Hankeling increases the mean human-normalized score from 0.463 to 0.530 across 26 games, an improvement of approximately 14%. Across the reported MuJoCo tasks, Q-Hankeling improves final return at 1M environment steps in seven of eight evaluated SAC/TD3 task-algorithm combinations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.