acceptodds
Under review as a conference paper at ICLR 2027

Beyond Self-Exploration: Stable Off-Policy Reinforcement Learning for LLMs

Abstract

Reinforcement learning (RL) has become a central paradigm for aligning LLMs, yet most existing approaches rely on on-policy self-exploration, which can lead to optimization stagnation when the policy is weak due to low-quality rollouts. Leveraging high-quality external trajectories offers a natural remedy, but this off-policy setting is notoriously unstable. Through an empirical analysis of PPO and GRPO, we identify unregulated importance ratios induced by off-policy distributions as a primary cause of this instability. Beyond instability, off-policy RL also suffers from attenuated gradients, as informative off-policy tokens typically receive low probability under the target policy. To address these challenges, we introduce Adaptive Projected Robust Optimization (APRO) to enable stable off-policy RL in LLM optimization. APRO constructs a tractable proxy for the unknown off-policy distribution and adaptively switches to a local policy anchor when proxy importance-ratio estimates become unreliable. In addition, APRO introduces discrepancy-aware advantage scaling to amplify informative learning signals that would otherwise be suppressed. We evaluate APRO on mathematical reasoning and instruction-following tasks in fully off-policy settings. Experimental results show that APRO achieves stable optimization and improves over the evaluated offline distillation and ratio-based off-policy RL baselines. Further experiments show that APRO can be used in mixed-policy training and can provide stronger initializations for subsequent on-policy RL, highlighting stable off-policy RL as a useful capability-transfer stage for weak target policies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.