acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Depth Scaling in Reinforcement Learning Through Horizon Generalization

Abstract

Scaling neural networks has driven major advances in supervised and self-supervised learning, yet the role of network depth in reinforcement learning remains unclear. Prior studies report limited gains from deeper networks, while a recent study finds substantial benefits only in self-supervised reinforcement learning. We reveal a missing piece of this puzzle: depth has a limited effect on in-distribution performance but substantially improves generalization to unseen goals. We study horizon generalization in locomotion, manipulation, and navigation tasks across self-supervised, off-policy, and on-policy reinforcement learning methods. In most settings, deeper policies maintain higher success rates as evaluation goals move farther beyond the training distribution, even when shallow and deep models perform similarly on training goals. Deeper networks also outperform wider, shallow networks with comparable parameter counts, showing that the gains are not due to model size alone. Mechanistic analyses further reveal depth-dependent differences in how policies represent and use relative goal information. Our results suggest that network depth primarily extends the range over which reinforcement learning policies generalize, and that standard in-distribution evaluations may underestimate its value.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.