acceptodds
Under review as a conference paper at ICLR 2027

Robust RLHF for LLM Alignment

Abstract

Reinforcement Learning with Human Feedback (RLHF) aligns large language models (LLMs) with human preferences, but its reliance on static, offline preference data leaves aligned models vulnerable to shifts in deployment contexts and user preferences. We introduce two distributionally robust (DR) alignment formulations under shifts measured by KL divergence, along with tractable PPO-based algorithms for each: JoRAL (joint context-preference shift robust alignment) and CoRAL, a simpler context-shift formulation in which the preference model remains unchanged. By directly connecting the log-linear policy class commonly used in theoretical analyses to modern LLMs, we prove that both CoRAL and JoRAL achieve sample complexity for the frozen-backbone, last-layer LLM policy class. Experiments on LLaMA-3.2-1B-Instruct show that JoRAL provides the most consistent and superior performance improvements across shifted evaluation environments, generally outperforming non-robust PPO and DPO and comparing favorably with existing robust-alignment methods, while CoRAL exhibits effective robustness to context shifts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.