acceptodds
Under review as a conference paper at ICLR 2027

What Off-Policy Corrections Keep: Gradient Direction and Realized Gains from Stale LLM Rollouts

Abstract

Language-model reinforcement learning increasingly learns from stale rollouts, produced by asynchronous systems or by long agentic interactions that are expensive to repeat, and continued training on them can erode earlier gains. We ask what learning signal different off-policy corrections extract from the same stale responses and when that signal yields realized improvement. We develop a theoretical framework for learning from stale data. It represents each correction through per-token coefficients that factor into advantage, weight, and gate; separates gradient magnitude from off-policy gradient alignment (OGA); and derives a local condition linking alignment and step size to realized improvement. Across mathematical reasoning and instruction following, correction rules produce markedly different directions even at low policy drift. Rules with identical ratio bounds reverse their majority OGA ordering across tasks, and near-equal scalar coefficient-mass totals can accompany large directional differences. OGA generally deteriorates as the policy moves away from the sampling policy, motivating drift as an observable state variable for reuse control. In held-out retrospective evaluation, we show that outcome-calibrated, drift-aware control of reuse retains gains that continued training would otherwise erase, without requiring fresh-gradient measurements.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.