Can You Trust Your Preference Models? Stackelberg Games for Alignment
Abstract
Language models are frequently aligned to human preferences via imperfect evaluators whose errors can be exploited during optimization. We train policies to search for comparisons on which an evaluator may reverse the underlying preference, providing candidates for additional supervision. Since human preferences are unavailable during policy training, we use disagreement among an ensemble of preference models as a proxy for prediction error. We formulate this search as a zero-sum Stackelberg game between two opposing policies, a Leader which tries to minimize disagreement, and a Follower which maximizes it. For a general class of finite contextual Stackelberg games, we establish equilibrium existence and, under regularization, uniqueness and a closed-form characterization. We introduce Stackelberg Mirror Descent–Ascent (SMIDA), a two-timescale algorithm with linear last-iterate convergence in the tabular setting, and develop a neural policy-gradient analogue. Experiments with a synthetic oracle support disagreement as an informative proxy for prediction error. On held out data the trained policies outperform reference policies and increase the fraction of comparisons that exceed a disagreement threshold and reverse the oracle’s preference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.