ISAR: Calibrating Historical Advantages for Efficient Trajectory Reuse
Abstract
Reinforcement learning improves language models through repeated generation and policy optimization, but autoregressive rollouts make training expensive. Reusing historical trajectories reduces generation demand, yet aggressive reuse can impair learning when their reward-relative signals become poorly matched to the evolving policy. We introduce ISAR, which recalibrates advantages by estimating the current policy's prompt-specific expected reward from fresh and historical trajectories. It combines full-trajectory, self-normalized importance weighting, candidate exclusion, and a circular replay buffer with a token-level clipped actor objective. Our analysis characterizes historical advantage evolution under policy improvement and shows how mixed-group averaging attenuates the response to that improvement. In search-agent training with Qwen3.5-9B, ISAR obtains 46.0 average exact match versus 44.5 for GRPO, using 60.5% fewer new tokens and 50.2% less total training time at the same update horizon. It further reduces these costs by 18.2% and 12.3% relative to the most token-efficient and fastest strong baselines, respectively. Full-parameter DeepSeek-V4-Flash training on operations research tasks reduces generation by 54.9% and time by 38.7%, with similar aggregate quality. Component controls show that calibration recovers quality lost under uncalibrated replay, supporting current-reference estimation as a practical component of efficient reuse.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.