acceptodds
Under review as a conference paper at ICLR 2027

Which Preferences to Train On? End-to-End Multi-Objective Alignment with an Adversarial Preference Distribution

Abstract

Aligning large language models (LLMs) with human values is important for safe, efficient, and beneficial AI deployment. However, human values are multifaceted: helpfulness, harmlessness and humor trade off against one another, and different users want different trade-offs. Multi-objective alignment (MOA) addresses this by training a policy that can provide any point of the Pareto front, but existing methods either train one model per preference, interpolate a few separately aligned experts post hoc, or train a single conditioned model without considering which preferences it should be trained on. Since the hard regions of the preference simplex depend on the objectives at hand, existing methods leave them under-trained and do not get the most out of a single model. Therefore, we propose MAESTRO (**M**ulti-objective **A**lignment via **E**nd-to-end **ST**eering and **R**obust **O**ptimization), which trains a single prompt-conditioned policy end-to-end with RL against an **adversarial** preference distribution: a Dirichlet distribution updated by online mirror descent toward the preferences the current policy serves worst, rather than on a fixed distribution. On HH-RLHF, BeaverTails and summarization, with up to three objectives, MAESTRO attains the best Pareto front on most tasks in a single training run, at the lowest training cost among all methods. The largest margins appear in the hard regions that a fixed preference distribution leaves under-trained, confirming that a single prompt-conditioned model is capable of covering the objective trade-offs on its own.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.