Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
Abstract
Reinforcement learning (RL) has emerged as a key paradigm for enhancing the reasoning capabilities of large language models. However, the high dimensionality of parameter updates makes RL training dynamics difficult to analyze, obscuring the mechanisms underlying these gains. In this work, we study reinforcement learning with verifiable rewards (RLVR) and use vector steering as an analytical tool to identify a low-dimensional effective manifold in activation space that is closely associated with RL-induced performance gains. We further uncover two key geometric properties of this manifold. (1) Effective Manifold Capacity: The capacity required to reproduce RL-induced gains can be very small, but it is not infinitely compressible. When the capacity is compressed to an extremely low level, the intervention dimensionality and the ability to express input-dependent corrections become key constraints, and this requirement further varies with injection depth. (2) Control Manifold Separation: Effective control directions lie predominantly in the low-variance complement of the activation principal subspace. Within the same task and base model, the learned geometry remains largely consistent across training configurations; across tasks, geometric alignment correlates with capability transfer. We conduct experiments on 5 LLMs and 6 tasks with verifiable rewards, and the results support the above properties. Based on these findings, we propose Alpha-Stabler, a plug-and-play training framework comprising two modules: a Predictor that monitors principal-subspace intrusion to provide early warnings of training collapse, and a Controller that removes the principal-subspace component of activation gradients during backpropagation while preserving their components in the orthogonal complement of the principal subspace. Experiments show that Alpha-Stabler can stabilize training for 2,000 steps and consistently enhance the gains brought by reinforcement learning. This work advances the understanding of RL from an activation-manifold perspective and offers practical insights for more robust post-training. Our code is available at: https://anonymous.4open.science/r/alpha-stabler.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.