SteerDPO: Steerable Multi-Objective Alignment via Implicit Reward Scalarization
Abstract
Aligning large language models (LLMs) with diverse, often conflicting human preferences is a fundamental challenge in building versatile, controllable AI assistants. While multi-objective alignment aims to address such trade-offs, existing approaches typically require costly online training or multiple explicit reward models, or rely on heuristic parameter-level aggregation with limited guarantees about the resulting preference trade-offs. We introduce SteerDPO, a single-stage, reward-model-free offline multi-objective alignment method that learns a single steerable policy directly from per-objective pairwise preference data. SteerDPO enables inference-time steering across trade-offs by explicitly conditioning on arbitrary preference weights, without additional training or online data collection. To this end, we derive the optimal policy for any preference weight as a normalized weighted geometric mean of the one-hot optimal policies under KL-regularized linear scalarization. SteerDPO uses this identity to train a single preference-conditioned policy through offline self-distillation. Two- and three-objective synthetic bandits and LLM experiments evaluate reward quality, steering, and front geometry separately. SteerDPO leads the compared methods on requested reward and steerability in the three-seed Llama-3.2-3B and single-seed Qwen3-8B comparisons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.