Beyond RLHF: A Theoretical Framework of LLM Alignment
Abstract
Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) to produce higher-quality responses. However, the standard training objective for RLHF lacks a learning-theoretic justification, and existing theories do not allow comparisons of the guarantees between various methods because different methods are often analyzed under different frameworks. Toward a unified framework of alignment, we ask under what assumptions we can justify existing algorithms or derive new ones. To this end, we reframe alignment as distribution learning from pairwise preferences, which makes a probabilistic assumption describing how preferences reveal information about the target LM. This leads us to derive three principled training objectives: preference maximum likelihood estimate, preference distillation, and reverse KL minimization. These can be viewed as corrections to existing objectives that, as we show, enjoy strong non-asymptotic convergence to the target LM under our framework. Furthermore, our theory explains, for the first time, the empirical finding that on-policy objectives (e.g., RLHF) typically outperform likelihood-style objectives (e.g., DPO). Our empirical results indicate that our algorithms are competitive with strong baselines across several tasks and models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.