acceptodds
Under review as a conference paper at ICLR 2027

Recovering RLVR from Failing Spurious Rewards with Subspace Regularization

Abstract

Reinforcement learning with verifiable rewards (RLVR) has driven large gains in LLM reasoning, yet it remains unclear how the reward signal shapes the resulting parameter update. Recent work reports a striking model-dependent behavior, where weak or spurious rewards can improve Qwen models while severely degrading other families such as Llama models. We investigate this failure through the geometry of reward-induced parameter updates and empirically find that the dominant output-side direction is substantially more reward-dependent than the corresponding input-side direction. This finding suggests a simple intervention: rather than replacing a failing reward, we can guide its parameter updates toward the direction induced by the true reward. Based on this insight, we introduce a subspace regularizer that guides weak-reward training using the dominant output-side direction obtained under true rewards. Across reward settings that otherwise fail, this guidance substantially recovers performance across mathematical reasoning benchmarks. While obtaining this direction from fully labeled training data would require extensive ground-truth supervision, we further show that it can be reliably estimated from only a small number of ground-truth examples, retaining much of the benefit of full-ground-truth guidance. Together, our results show that the update direction induced by a reward provides an actionable signal for recovering RLVR when weak or spurious rewards fail.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.