acceptodds
Under review as a conference paper at ICLR 2027

Future-Aligned Reinforcement Learning for Asynchronous Language Model Post-Training

Abstract

Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, at the cost of staleness: a rollout is generated by a behavior policy that is one optimizer update behind the target policy by the time it is consumed. Most existing methods address this staleness reactively, for example through importance-sampling and clipping schemes or distribution-matched replay frameworks, repairing stale data after it has been produced. We introduce Future-Aligned Reinforcement Learning (FARL), a proactive method that, at rollout-dispatch time, forecasts the consumption-time policy (the learner one update ahead) and generates the entire trajectory from this single forecasted, frozen worker policy. The forecast extrapolates the learner's next parameter displacement from the most recent updates using a consistency-constrained multi-lag predictor. We show that the consistency constraint cancels the first-order constant-drift component of policy lag: under fixed-step smooth learner dynamics this yields an asymptotic upper bound on the forecast residual against for stale generation, where is the learner's step size, and correspondingly tightens the local bound on the variance of the exact importance weights from to . For finite-step stochastic training we give a separate coefficient-aware data-dependent bound and measure the realized prequential forecast error directly. Because FARL acts on the behavior policy before generation rather than repairing data afterward, it is a drop-in plug-in module orthogonal to existing reactive correctors, and across a thorough set of experiments it improves the average accuracy of every one of them. Code is available at https://anonymous.4open.science/r/farl-anonymous-0442.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.