acceptodds
Under review as a conference paper at ICLR 2027

Robust Preference Alignment under Shifting Preference Distributions

Abstract

Current preference-based alignment techniques for large language models typically learn from fixed preference labels that order candidate responses for a given prompt. However, human judgments are multidimensional and context-dependent, so different but valid evaluation priorities may reorder the same responses. Standard objectives treat the observed order as a fixed target and therefore do not account for this variation. Consequently, when deployment priorities differ from those represented during training, a policy fitted to the training labels may not remain aligned. Existing robust methods mainly interpret deviations from observed labels as unreliable feedback at sample-level. We instead study robustness to conditional ranking shifts and propose Kendall Robust Preference Optimization (KRPO). KRPO captures the empirical structure of these shifts through a utility-gap-weighted Kendall ambiguity set, making reversals between similarly scored responses less costly than those between clearly separated responses. An entropy-regularized adversary then emphasizes difficult but plausible alternative rankings. Because KRPO changes only the ranking supplied to the base loss, it applies to both pairwise and listwise objectives. Experiments with seven baseline objectives on Llama-3.1-8B and Qwen3-1.7B show that KRPO consistently improves cross-benchmark generalization and surpasses sample-level robust baselines on identical, unperturbed data. KRPO thus shows that robustness to shifting evaluation priorities is achieved by modeling where a ranking is likely to change and it offers a general mechanism for extending existing preference objectives to criteria absent from their supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.