acceptodds
Under review as a conference paper at ICLR 2027

LLM Alignment with Biased Proxy Preferences under Heterogeneous Utilities

Abstract

Aligning a language model in a specialized domain such as medicine raises the question of whose judgment should count. Expert comparisons are on target but scarce, a general population's are plentiful but systematically biased, and pooling the two learns a utility between the two populations. We ask when the crowd's comparisons still help, which way of using them helps most, and why. Under heterogeneous utilities, we propose a two-headed estimator (THE) of the direction of the experts' utility center from a small gold and a large proxy sample, which fits one utility per population with the sign loss and couples the two through a penalty on their difference; re-weighting the pooled samples is its infinite-penalty limit. We prove that proxy comparisons speed up estimation while their bias is below the statistical error of pooling, and that once the bias exceeds the gold-only error the gain of any re-weighting over gold-only vanishes as the bias grows. For controlled comparisons, which vary one group of criteria at a time, with few contentious criteria and precisely estimated consensus criteria, a variant of THE has a smaller asymptotic error than every re-weighting. With a smoothed loss, a finite penalty improves on the best re-weighting whenever the bias lies, on average, in more curved directions of the gold loss than the error of the best re-weighting. For a bias sparse in of features, THE's two-step special case with an penalty reduces the gold sample size that our sign-loss bounds require by a factor of order , under local conditions. In simulations with the exact sign loss, THE improves on gold-only by 15% at a bias twice the gold-only error, where the best re-weighting improves by 3%; on real answer embeddings with simulated preferences, it improves on the best re-weighting by up to 20% when the bias lies along high-variance features.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.