Last-Iterate Convergence of Optimistic Primal–Dual Methods for KL-Regularized Constrained RLHF
Abstract
Reinforcement Learning from Human Feedback (RLHF) plays a significant role in aligning Large Language Models (LLMs) with human preferences. RLHF with expected reward constraints can be formulated as a KL-regularized saddle-point problem, in which simultaneous primal–dual updates can exhibit oscillatory dynamics. We study an optimistic predictor–corrector method that applies optimism to both the policy and the constraint multipliers. In the distributional policy space, we prove geometric last-iterate convergence to the regularized saddle-point set. For parameterized policies, we give an inexact-oracle extension whose residual separates primal solver error and comparator coverage under an additional uniform relative-density bound. Experiments with a PPO-style surrogate on PKU-SafeRLHF-30K illustrate reduced oscillation and improved final-iterate robustness over the implemented PPO-Lagrangian baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.