What Off-Policy Corrections Keep: Gradient Direction and Realized Gains from Stale LLM Rollouts
Abstract
Language-model reinforcement learning increasingly learns from stale rollouts, produced by asynchronous systems or by long agentic interactions that are expensive to repeat, and continued training on them can erode earlier gains. We ask what learning signal different off-policy corrections extract from the same stale responses and when that signal yields realized improvement. We develop a theoretical framework for learning from stale data. It represents each correction through per-token coefficients that factor into advantage, weight, and gate; separates gradient magnitude from off-policy gradient alignment (OGA); and derives a local condition linking alignment and step size to realized improvement. Across mathematical reasoning and instruction following, correction rules produce markedly different directions even at low policy drift. Rules with identical ratio bounds reverse their majority OGA ordering across tasks, and near-equal scalar coefficient-mass totals can accompany large directional differences. OGA generally deteriorates as the policy moves away from the sampling policy, motivating drift as an observable state variable for reuse control. In held-out retrospective evaluation, we show that outcome-calibrated, drift-aware control of reuse retains gains that continued training would otherwise erase, without requiring fresh-gradient measurements.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.