Axiom Satisfiability of Linear Rewards in Alignment
Abstract
Learning from human preference data is the dominant route to aligning language models with human values. In the linear social choice model, where rewards are linear in a fixed feature representation of prompt-response pairs, Ge et al. (2024) show that fitting such a reward by minimizing any non-decreasing convex loss, including Bradley-Terry-Luce (BTL), fails Pareto Optimality (PO) and Pairwise Majority Consistency (PMC). Moreover, no rule reading only the majority relation can satisfy PO once the output is required to be linearly induced. We ask what it costs to enforce these axioms anyway. To this end, we relax the linear model to allow per-candidate additive slack. We compute the relaxed linear reward with the smallest total slack that satisfies the axioms with a margin , the minimum required difference between two reward values. Our solution satisfies the axioms under no assumptions about the voters or how comparisons were collected. We bound the optimal total slack by when is at most for candidates. Furthermore, we exhibit an instance that forces this bound, concluding that the rate is tight up to constants. In practice, the number of candidates far exceeds the feature dimension, and only a linear reward can be evaluated on unseen responses. We therefore introduce a new method that charges the linear part for each comparison it gets wrong. It simultaneously minimizes the total slack and the number of violations, with a parameter that trades off between them. We show that the total slack is monotone but saturating in : raising it reduces the violations of the deployed linear reward and increases the slack, yet the slack stays below , where and are the diameter and dimension of the features, respectively. Experiments on both synthetic and real-life preference data with per-annotator labels corroborate our theoretical result and show that the linear reward output by our method beats linear BTL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.