acceptodds
Under review as a conference paper at ICLR 2027

COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning

Abstract

Asynchronous RL improves throughput in large language model post-training by decoupling rollout generation from optimization, at the cost of training on stale trajectories. Existing methods primarily address token-level policy mismatch through importance-ratio control in the actor objective, providing policy-side correction. We show that this alone is insufficient because advantage estimates inherit mismatch from behavior-policy continuations, a phenomenon we term advantage staleness. This motivates a complementary advantage-side correction, raising the question: how do the two correction channels interact, and how should their strengths be coordinated? We derive exact bias and variance decompositions for a general two-channel actor update, revealing a nonseparable coupling between policy-weight and advantage-estimation errors: their interaction induces multiplicative bias terms, while squared policy weights amplify advantage uncertainty in gradient variance. This coupling motivates a testable hypothesis that the preferred settings of the two correction channels should be coordinated rather than chosen independently. Guided by this analysis, we introduce Coupled Off-Policy Correction (COPC), an actor–critic method combining token-level ratio masking with a dedicated advantage-side correction based on two-sided clipped-ratio weighting of TD residuals for return and advantage estimation. Joint parameter sweeps across staleness levels provide empirical support for this hypothesis: the effect of varying one correction parameter depends on, and can even reverse with, the setting of the other. COPC achieves the highest reported performance on both tool-integrated mathematical reasoning and search, outperforming the strongest reported asynchronous baseline in each setting. COPC further exhibits strong practical robustness, with a broad high-performing parameter region and substantially improved training stability. In particular, in the search setting, it remains stable throughout training, while most evaluated asynchronous baselines exhibit late-stage performance collapse. These performance gains persist at 64-step policy staleness, while COPC incurs minimal step-time overhead over asynchronous PPO and retains a step-time speedup over synchronous PPO.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.