acceptodds
Under review as a conference paper at ICLR 2027

ImPlay: Importance-Weighted Experience Replay for Efficient Post-Training

Abstract

Group Relative Policy Optimization (GRPO) has become a standard algorithm for post-training large language models (LLMs). However, each update depends on many on-policy rollouts, making generation the dominant training cost. Experience replay can reduce this cost by reusing past rollouts, but GRPO's group-relative advantages become stale as the policy evolves. We first analyze off-policiness in GRPO with experience replay and find that advantages can become substantially stale after only 10–20 optimizer steps, while sample age is only weakly indicative of this staleness. Motivated by this analysis, we propose ImPlay, an experience replay method for GRPO that introduces two mechanisms to address policy drift: importance-weighted advantage re-estimation, which uses importance sampling to re-estimate replayed trajectories’ group-relative advantages under the current policy; and drift-aware eviction, which uses the same signal to identify and discard trajectories that have become too off-policy for reliable reuse. Evaluated across six mathematical reasoning benchmarks, ImPlay matches GRPO's performance under the same training setup while requiring up to 47% fewer generated rollouts and 12% faster training. At a matched generation budget, ImPlay outperforms ExGRPO, the state-of-the-art experience-replay method for GRPO, by up to 3.0 percentage points on the more challenging benchmarks. Ablations show that drift-aware eviction maintains a replay buffer in which importance-weighted advantage re-estimation remains effective.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.