Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives zero learning signal. Providing privileged guidance during training, such as solution prefixes, overcomes this learning cliff by steering the model towards rewarding generations. We observe that these guided generations are fundamentally off-context, i.e., there is a mismatch between the inference and training distributions. Our key contribution is to address this mismatch, ensuring the model ultimately learns to solve problems without guidance. To this end, we introduce Off-Context GRPO (OC-GRPO): a minimally modified variant of GRPO with an importance-corrected objective. OC-GRPO explicitly adjusts for context mismatch, cleanly separating guidance (where to explore) from the objective (what to optimize). By provably recovering an unbiased gradient estimator for the original unguided objective, OC-GRPO prevents instabilities during training. Empirically, our algorithm achieves a 3.9% absolute improvement (13.8% relative gain) over vanilla GRPO on average across standard mathematical reasoning benchmarks without adding any extra computational complexity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.