Multi-Objective LLM Alignment using Multivariate Stochastic Dominance
Abstract
Though Large Language Models (LLMs) are increasingly performant on a wide range of tasks, multi-objective alignment with a variety of evaluative signals is a key challenge. Existing methods typically optimize scalar preference signals, seek to balance competing objectives, or learn policies spanning different trade-offs, without explicitly targeting joint improvement over a reference policy. In this paper, we introduce Stochastically Dominant Multi-Objective Alignment (SDMA), a framework for multi-dimensional LLM alignment without committing to particular trade-offs between evaluative dimensions. Instead, we employ first-order multivariate stochastic dominance to guide improved performance for all monotonic functions of these different dimensions. We demonstrate balanced improvement in our results for multiple dimensions, including subsets of helpfulness, harmlessness, humor, honesty, and instruction-following.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.