acceptodds
Under review as a conference paper at ICLR 2027

Inter-Iteration Policy Improvement for LLM Reinforcement Learning

Abstract

Reinforcement learning is central to post-training LLMs for reasoning and agentic tasks. However, existing methods optimize rewards, advantages, or distillation targets constructed from the current batch without using subsequent performance observations to inform later optimization. This open-loop structure leaves recent policy progress outside the learning signal and makes optimization sensitive to finite sampling, generation stochasticity, and noisy feedback. We introduce Policy Improvement Reinforcement Learning (PIRL), a trajectory-level formulation that makes inter-iteration policy improvement an explicit optimization signal while retaining cumulative equivalence to terminal task performance. Building on PIRL, we propose Policy Improvement Policy Optimization (PIPO), a plug-in framework that closes the loop around policy updates. At each iteration, PIPO compares current empirical performance with a sliding-window historical anchor and uses the resulting progress signal to retrospectively modulate the base algorithm's local attribution. We establish local alignment with isolated base-step improvement under explicit trend-coherence conditions and characterize signed retrospective replay for group-relative optimization. Across two model scales and mathematical reasoning, code, tool-use, and self-distillation tasks, PIPO consistently improves PPO, GRPO, GSPO, DAPO, and SDPO under matched training and evaluation settings across all reported experiments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.