Three Gains, One Loop: Diagnosing and Repairing the Bellman Bootstrap in High Dimensions
Abstract
Model-free reinforcement learning is widely believed to require world models or complex architectural innovations to scale to high-dimensional continuous control. We show that this difficulty can stem from a core structural instability: in high dimensions, the Bellman bootstrap forms a positive-feedback loop that can amplify noise faster than it accumulates signal. As the agent improves, its actions move into poorly covered regions, generating out-of-distribution temporal-difference errors that can corrupt the value landscape needed for continued learning. We identify three gains governing this feedback loop, corresponding to target signal, critic-gradient geometry, and error sensitivity, and show how each can be compressed by an architecture-preserving intervention: multi-step returns, LayerNorm, and Huber loss, respectively. We term this repair Gain-Compressed Repair (GCR). While the individual components are established techniques, the gain decomposition explains their complementary roles and predicts distinct failure signatures when individual gains are left uncontrolled. Together, these modifications add negligible parameter overhead yet enable standard TD3 and SAC to achieve nearly 800 reward on Dog-Run (38-D actions) within 2M steps and strong performance on 61-dimensional H1Hand tasks. Across 12 primary high-dimensional benchmarks spanning Dog, Humanoid, and H1Hand, GCR remains competitive with strong model-based baselines, suggesting that Bellman-loop instability is a key bottleneck for model-free learning in the studied state-based settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.