acceptodds
Under review as a conference paper at ICLR 2027

Can You Trust Your Preference Models? Stackelberg Games for Alignment

Abstract

Language models are frequently aligned to human preferences via imperfect evaluators whose errors can be exploited during optimization. We train policies to search for comparisons on which an evaluator may reverse the underlying preference, providing candidates for additional supervision. Since human preferences are unavailable during policy training, we use disagreement among an ensemble of preference models as a proxy for prediction error. We formulate this search as a zero-sum Stackelberg game between two opposing policies, a Leader which tries to minimize disagreement, and a Follower which maximizes it. For a general class of finite contextual Stackelberg games, we establish equilibrium existence and, under regularization, uniqueness and a closed-form characterization. We introduce Stackelberg Mirror Descent–Ascent (SMIDA), a two-timescale algorithm with linear last-iterate convergence in the tabular setting, and develop a neural policy-gradient analogue. Experiments with a synthetic oracle support disagreement as an informative proxy for prediction error. On held out data the trained policies outperform reference policies and increase the fraction of comparisons that exceed a disagreement threshold and reverse the oracle’s preference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.