Inter-Iteration Policy Improvement for LLM Reinforcement Learning
Abstract
Reinforcement learning is central to post-training LLMs for reasoning and agentic tasks. However, existing methods optimize rewards, advantages, or distillation targets constructed from the current batch without using subsequent performance observations to inform later optimization. This open-loop structure leaves recent policy progress outside the learning signal and makes optimization sensitive to finite sampling, generation stochasticity, and noisy feedback. We introduce Policy Improvement Reinforcement Learning (PIRL), a trajectory-level formulation that makes inter-iteration policy improvement an explicit optimization signal while retaining cumulative equivalence to terminal task performance. Building on PIRL, we propose Policy Improvement Policy Optimization (PIPO), a plug-in framework that closes the loop around policy updates. At each iteration, PIPO compares current empirical performance with a sliding-window historical anchor and uses the resulting progress signal to retrospectively modulate the base algorithm's local attribution. We establish local alignment with isolated base-step improvement under explicit trend-coherence conditions and characterize signed retrospective replay for group-relative optimization. Across two model scales and mathematical reasoning, code, tool-use, and self-distillation tasks, PIPO consistently improves PPO, GRPO, GSPO, DAPO, and SDPO under matched training and evaluation settings across all reported experiments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.