Uncovering Cross-Objective Interference in Multi-Objective Alignment
Abstract
We study a persistent failure mode in multi-objective alignment for large language models (LLMs), in which scalarized training improves only a subset of objectives while the others degrade. We formalize this phenomenon as cross-objective interference and conduct, to our knowledge, the first systematic study of nine scalarization algorithms, finding it pervasive yet strongly model-dependent. To explain interference, we derive a local covariance law in which an objective rises or falls with the sign of its covariance with the scalarized score. We extend this analysis to clipped surrogate objectives used in modern alignment, demonstrating that the covariance law remains valid under mild conditions despite clipping. Building on this analysis, we propose COVariance-floor Enforced Reweighting (COVER), a plug-and-play controller that raises an objective's weight only when its reward's covariance with the clipped advantage weight falls below a target. Through extensive experiments, we show that COVER effectively mitigates cross-objective interference while matching linear scalarization when objectives already co-improve. Finally, we complement these local improvement conditions with a global convergence analysis that establishes when non-convex scalarized optimization satisfies the Polyak–ojasiewicz condition and how cross-objective interference depends on model geometry.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.