Recovering RLVR from Failing Spurious Rewards with Subspace Regularization
Abstract
Reinforcement learning with verifiable rewards (RLVR) has driven large gains in LLM reasoning, yet it remains unclear how the reward signal shapes the resulting parameter update. Recent work reports a striking model-dependent behavior, where weak or spurious rewards can improve Qwen models while severely degrading other families such as Llama models. We investigate this failure through the geometry of reward-induced parameter updates and empirically find that the dominant output-side direction is substantially more reward-dependent than the corresponding input-side direction. This finding suggests a simple intervention: rather than replacing a failing reward, we can guide its parameter updates toward the direction induced by the true reward. Based on this insight, we introduce a subspace regularizer that guides weak-reward training using the dominant output-side direction obtained under true rewards. Across reward settings that otherwise fail, this guidance substantially recovers performance across mathematical reasoning benchmarks. While obtaining this direction from fully labeled training data would require extensive ground-truth supervision, we further show that it can be reliably estimated from only a small number of ground-truth examples, retaining much of the benefit of full-ground-truth guidance. Together, our results show that the update direction induced by a reward provides an actionable signal for recovering RLVR when weak or spurious rewards fail.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.