acceptodds
Under review as a conference paper at ICLR 2027

Multi-Objective LLM Alignment using Multivariate Stochastic Dominance

Abstract

Though Large Language Models (LLMs) are increasingly performant on a wide range of tasks, multi-objective alignment with a variety of evaluative signals is a key challenge. Existing methods typically optimize scalar preference signals, seek to balance competing objectives, or learn policies spanning different trade-offs, without explicitly targeting joint improvement over a reference policy. In this paper, we introduce Stochastically Dominant Multi-Objective Alignment (SDMA), a framework for multi-dimensional LLM alignment without committing to particular trade-offs between evaluative dimensions. Instead, we employ first-order multivariate stochastic dominance to guide improved performance for all monotonic functions of these different dimensions. We demonstrate balanced improvement in our results for multiple dimensions, including subsets of helpfulness, harmlessness, humor, honesty, and instruction-following.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.