Preference-aware Multi-objective Alignment
Abstract
Alignment is an important component of post-training for modern foundation models and often involves multiple conflicting objectives with heterogeneous preferences specified by stakeholders. Failing to account for these conflicts and preferences can lead to undesirable trade-offs and unstable optimization. In practice, however, objective-level preferences may be only partially specified, e.g., through qualitative orderings among objectives rather than precise numerical weights. To address this challenge, we propose a preference-aware multi-objective alignment framework that models preferences over objectives through an admissible preference set, enabling a unified representation of both cardinal and ordinal preference information, including partial order constraints. Building on this framework, we develop PrEference-Consistent Cardinal-Ordinal Alignment (PECCO), an algorithm that mitigates gradient conflicts among objectives while preserving the specified preferences throughout optimization. We further establish convergence to a preference-aligned Pareto stationary point. Experiments on standard benchmarks show that PECCO achieves performance comparable to existing multi-objective preference-alignment methods while providing greater flexibility in accommodating diverse forms of objective-level preferences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.