acceptodds
Under review as a conference paper at ICLR 2027

Beyond RLHF: A Theoretical Framework of LLM Alignment

Abstract

Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) to produce higher-quality responses. However, the standard training objective for RLHF lacks a learning-theoretic justification, and existing theories do not allow comparisons of the guarantees between various methods because different methods are often analyzed under different frameworks. Toward a unified framework of alignment, we ask under what assumptions we can justify existing algorithms or derive new ones. To this end, we reframe alignment as distribution learning from pairwise preferences, which makes a probabilistic assumption describing how preferences reveal information about the target LM. This leads us to derive three principled training objectives: preference maximum likelihood estimate, preference distillation, and reverse KL minimization. These can be viewed as corrections to existing objectives that, as we show, enjoy strong non-asymptotic convergence to the target LM under our framework. Furthermore, our theory explains, for the first time, the empirical finding that on-policy objectives (e.g., RLHF) typically outperform likelihood-style objectives (e.g., DPO). Our empirical results indicate that our algorithms are competitive with strong baselines across several tasks and models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.