acceptodds
Under review as a conference paper at ICLR 2027

Follow the Improvement: Guaranteed Policy Optimization by Tracking Proximal Ascent Targets

Abstract

Modern policy optimization approaches are mainly built on two theoretical paradigms: constrained optimization and tracking a target policy. Constrained methods (PPO, TRPO) maximize a surrogate objective within a trust region, but their monotonic improvement guarantees hold only for an idealized procedure, not for its practical implementation. Target-tracking methods (SAC, AWR, MPO) construct a target with improved performance and approach it by minimizing a regression loss; a small tracking loss however does not guarantee improvement. We propose *Follow-the-Improvement RL (FIRL)*, a framework that *unifies* these two main classes of policy optimization paradigms. First, FIRL defines a general proximal ascent target, , where aligns with the improvement direction and has constrained scale. Second, FIRL characterizes a class of *tracking loss functions* that can bound the performance gap between policies. Then we show that, with a constrained step size, minimizing idealized tracking losses toward the target , evaluated under the data distribution from , is sufficient to *guarantee performance improvement* over . The framework unifies PPO, SPO, and on-policy variants of AWR and MPO, while identifying *new promising instances* motivated by the theory: a clipped-linear target paired with advantage-weighted total variation/hinge (ATV/AH) losses. Across MuJoCo, IsaacLab, Atari, and LLM fine-tuning on mathematical reasoning, we evaluate different design combinations, show that the proposed instances perform competitively, and identify more effective target–loss combinations for each task category.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.