Just Stack It: Deeper Reinforcement Learning with Near-Isometric Layers
Abstract
Deep networks have driven progress in vision and language, yet reinforcement learning (RL) agents still rely on networks a few layers deep, because adding layers degrades their performance. We show that this failure results from a loss of conditioning: during training, the Jacobian of a deep network's hidden layers moves away from an isometry, its features lose rank, and the deviations of the layers multiply with depth. Under the non-stationary, bootstrapped objective of RL, this degradation impairs the fitting of later targets. The distance of the Jacobian from an isometry and the rank of the features track performance across architectures and optimizers. We propose a simple residual layer that keeps deep networks close to an isometry. With it, deep networks train stably across value-based, actor-critic and goal-conditioned algorithms and optimizers, and their performance improves or remains stable with depth: they outperform the standard shallow networks and recent deep architectures at matched depth and parameters.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.