acceptodds
Under review as a conference paper at ICLR 2027

Tail-Aware Direct Preference Optimization

Abstract

The geometry of policy regularization is a central design choice in LLM alignment. KL-regularized RLHF and DPO are attractive because they yield exponential tilting and a simple log-ratio preference loss, but this geometry can amplify low-coverage reward-model errors and cause severe overoptimization. We propose Tail-Aware Direct Preference Optimization (TDPO), a direct preference method derived from R\'enyi-regularized RLHF. TDPO derives an exact direct-preference loss on the active constraint, whose Bradley–Terry score difference is a difference of density-ratio powers rather than the KL-induced log-ratio. The order controls statistical coverage, tail sensitivity, and anti-hacking strength as a hyperparameter of the preference law. We prove finite-sample bounds for the unregularized optimality gap and regularized sub-optimality gap under a Bradley–Terry preference model. A two-environment lower bound matches the latter's sample-size exponent for arbitrary policy-valued estimators. We further characterize the induced geometry: locally, R\'enyi regularization rescales the Fisher geometry of KL, while globally it provides polynomial rare-set control. A TL;DR study evaluates density-ratio tails and low-reference-coverage mass across DPO, ChiPO, and clipped/capped baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.