Reward-Gradient Aligned Low-Rank Adaptation for Post-Training
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has become a prominent post-training paradigm for enhancing the reasoning capabilities of large language models. Low-Rank Adaptation (LoRA) offers a natural way to reduce its training cost. Intuitively, training both low-rank factors should allow LoRA to adapt along the directions that matter for the reward. However, we observe that the input-side factor of LoRA barely changes during RLVR. As a result, the input subspace remains close to its random initialization throughout training. We prove that under standard initialization this subspace rotates only at second order in the cumulative gradient, one order slower than the output-side factor. The random initialization therefore persistently constrains subsequent optimization. Motivated by this finding, we propose REward-Gradient Aligned Low-Rank Adaptation (REAL), which aligns the input subspace with task-specific reward gradients at initialization. REAL contrasts the gradients of correct and incorrect rollouts for each prompt and selects the subspace that retains the most energy of these contrasts across prompts. Since exact selection requires a full weight gradient for every prompt, REAL estimates the subspace by two-sided random compression with memory linear in the layer width. Extensive experiments on the Qwen3 series show that REAL outperforms seven PEFT baselines and exceeds LoRA by up to 5.90% in average accuracy. With a length-aware selection reward, REAL also shortens the responses of distilled reasoning models by up to 44.9% while improving their accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.