Estimation Debiasing for Asynchronous Policy Optimization
Abstract
In reinforcement learning for large language models, asynchronous pipelines sample a batch of rollouts once and use them for several updates. This reduces generation overhead but degrades performance due to rollout staleness: the behavior policy gradually diverges from the target policy, introducing off-policy errors that bias the estimated objective. Existing methods control this error with PPO-style clipping by limiting the policy shift. However, our derivation shows that this restriction also limits objective improvement, making clipping alone insufficient. In addition to clipping, the bias should be directly corrected in the objective estimator. We propose Bias-Corrected Yielded Policy Optimization (BYPO), which constructs a more accurate objective estimator through sequence-level importance weighting and second-order Taylor correction, adaptively switches between them to control variance, and falls back to a first-order approximation when both corrections become excessively large. Experiments on mathematical reasoning and preference-alignment tasks show consistent improvements in asynchronous training. Applied to DAPO, BYPO mitigates 83.70% of the performance degradation caused by rollout staleness, with negligible runtime overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.