Representational Manifold Hypothesis Enables Sample-efficient Online Contrastive Reinforcement Learning
Abstract
General-purpose mechanisms from supervised deep learning, such as residual blocks, normalization, and depth, have recently been shown to transfer to online reinforcement learning (RL) and to improve its sample efficiency. One further mechanism, the manifold hypothesis, has been observed in RL but used only architecturally, through bottlenecks whose fixed rank must be chosen before training. We instead make the representation low-dimensional by learning. Building on neural regression collapse, we propose NRC2Reg, a regularizer that adds no parameters and pulls intermediate features toward the row space of the network's own readout, with a single coefficient and no rank to tune. On Contrastive RL, the features settle near task intrinsic dimension at effective rank ten of a 64-dimensional budget, a regime no swept hard bottleneck reaches, while NRC2Reg improves sample efficiency in all seven environment and depth settings and final performance in six, with the largest gains on maze tasks where the unregularized agent barely learns.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.