acceptodds
Under review as a conference paper at ICLR 2027

Last-Iterate Convergence of Optimistic Primal–Dual Methods for KL-Regularized Constrained RLHF

Abstract

Reinforcement Learning from Human Feedback (RLHF) plays a significant role in aligning Large Language Models (LLMs) with human preferences. RLHF with expected reward constraints can be formulated as a KL-regularized saddle-point problem, in which simultaneous primal–dual updates can exhibit oscillatory dynamics. We study an optimistic predictor–corrector method that applies optimism to both the policy and the constraint multipliers. In the distributional policy space, we prove geometric last-iterate convergence to the regularized saddle-point set. For parameterized policies, we give an inexact-oracle extension whose residual separates primal solver error and comparator coverage under an additional uniform relative-density bound. Experiments with a PPO-style surrogate on PKU-SafeRLHF-30K illustrate reduced oscillation and improved final-iterate robustness over the implemented PPO-Lagrangian baseline.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.