Analyzing WEight-Space Orientation for Measuring Entropy and Accuracy in RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) can improve accuracy while reducing predictive entropy.We show these two effects are independent; their concomitance depends on a particular direction of the RLVR update and the entropy decrease involves token specific changes. Using two model lineages OLMo-3 and Tulu-3/Llama and their checkpoints, we perturb their weights and output scores. We find the components of RLVR update define a low dimensional subspace. We compare the RLVR weight update with randomly rotated versions in that subspace with the same singular values, and update size. The learned direction sharply lowers entropy, while the random directions have little effect. This difference persists across the tested truncation ranks and on model-generated text. We then use RLVR to select 30 strongly changed weight matrices, but search for alternative changes to those matrices around the supervised fine-tuned (SFT) model. We show that a similar measured accuracy can occur without entropy collapse and that entropy loss occurs with accuracy decreases in other points of the subspace.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.