acceptodds
Under review as a conference paper at ICLR 2027

Groupwise Distortion Guarantees for Preference-Based Alignment

Abstract

Preference-based alignment methods such as RLHF and NLHF aggregate ordinal pairwise comparisons to learn an LLM policy, but a natural goal is maximizing (average cardinal utility), which comparisons do not identify. Gölz, Haghtalab, and Yang (GHY) measure the gap by : the worst-case ratio between the welfare of the best fixed lottery (distribution over responses) and of the learned lottery. They show NLHF is optimal when every user receives the same lottery. Account-based LLMs, however, have information about their users and can serve different lotteries to different people. We give an efficient algorithm, GLHF, that learns a single group-conditioned policy from one comparison per user. This policy attains GHY's optimal distortion bound simultaneously on every group in a prespecified, possibly overlapping collection, with sample complexity growing logarithmically in the number of groups and inversely with the smallest group mass. When the members of a group share a feasible favorite response, that group's distortion approaches one. On human coffee ratings and synthetic LLM-generated ratings, GLHF lowers distortion in every group relative to NLHF and substantially reduces worst-group distortion relative to group-agnostic baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.