acceptodds
Under review as a conference paper at ICLR 2027

A Unified Random-Utility View of Heterogeneity in Direct Preference Optimization

Abstract

Direct Preference Optimization (DPO) assumes a single comparison scale, although in practice some response pairs are easy to distinguish while others remain ambiguous at similar utility gaps. Recent DPO-like variants increasingly introduce pair-adaptive temperatures, margins, confidence signals, or latent structure to address heterogeneity, yet similar language is often used for methods that model different forms of heterogeneity at different stages of preference modeling. To disentangle these differences, we present a unified random-utility view of heterogeneity in DPO and use it to organize DPO-like methods into a branch-level taxonomy. This taxonomy isolates a statistically distinct DPO-side error-layer branch, comparison-specific observation-scale heterogeneity, whose minimal instantiation is Heteroscedastic Direct Preference Optimization (H-DPO), a heteroscedastic-logit generalization of DPO that preserves the standard utility-gap parameterization while introducing comparison-specific scale inside the logit. In this branch, homoscedastic DPO is semantically misspecified: to match the correct preference probabilities, its score must absorb comparison-specific scale, converging to a scale-distorted quantity rather than the latent utility gap itself. Under restricted smooth families, the corresponding score-scale decomposition is locally identifiable up to global scaling from binary preference data. Synthetic experiments confirm the predicted score distortion, and real-world benchmarks provide empirical support for the comparison-scale formulation and its preference-discrimination benefits.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.