Robust RLHF for LLM Alignment
Abstract
Reinforcement Learning with Human Feedback (RLHF) aligns large language models (LLMs) with human preferences, but its reliance on static, offline preference data leaves aligned models vulnerable to shifts in deployment contexts and user preferences. We introduce two distributionally robust (DR) alignment formulations under shifts measured by KL divergence, along with tractable PPO-based algorithms for each: JoRAL (joint context-preference shift robust alignment) and CoRAL, a simpler context-shift formulation in which the preference model remains unchanged. By directly connecting the log-linear policy class commonly used in theoretical analyses to modern LLMs, we prove that both CoRAL and JoRAL achieve sample complexity for the frozen-backbone, last-layer LLM policy class. Experiments on LLaMA-3.2-1B-Instruct show that JoRAL provides the most consistent and superior performance improvements across shifted evaluation environments, generally outperforming non-robust PPO and DPO and comparing favorably with existing robust-alignment methods, while CoRAL exhibits effective robustness to context shifts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.