acceptodds
Under review as a conference paper at ICLR 2027

MOMoE: Instantiating Mixture-of-Experts for Multi-Objective LLM Alignment

Abstract

Aligning Large Language Models (LLMs) for diverse users and contexts requires balancing objectives such as helpfulness and harmlessness, yet no single trade-off between them suits all. Existing methods steer a single model to any user-specified preference at inference time, typically by composing objective-specific models through merging coefficients. However, they do not consistently surpass MORLHF, which trains a separate policy per preference. We trace this limitation to these methods adopting only the routing interface of a Mixture-of-Experts (MoE), with frozen, objective-specific experts and unregularized routing. We thus propose Multi-Objective Mixture-of-Experts (MOMoE) through four design principles: (1) joint optimization trains experts and router together rather than the router alone over frozen experts; (2) free experts are added with no objective assigned, reaching trade-offs the objective-specific ones cannot; (3) blend selection chooses each objective-specific expert by its synergy with the others rather than by its own objective; and (4) coverage regularization adds preference-conditioned losses that keep experts reachable and distinct. Extensive experiments show MOMoE outperforms state-of-the-art methods, attaining a broader Pareto front and higher utility (over 20% higher linear utility than MORLHF). We validate each principle through ablation and confirm gains through human evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.