Gap-Adaptive Per-Objective Regularization against Differential Overoptimization in Multi-Objective Alignment
Abstract
Multi-objective alignment scalarizes heterogeneous rewards into one objective optimized under a shared KL budget; as optimization strengthens, the policy overfits its reward proxies, and the resulting overoptimization is not distributed evenly across objectives. Existing methods either apply one global intervention (a single KL budget or cap) or only steer the direction of optimization, without adapting each objective’s budget to a measured overoptimization gap. We formalize the answer in a minimal model of proxy and gold and name the pattern Differential Overoptimization: damage concentrates on the objective with the largest effective hackability, and each objective’s collapse point follows a closed-form allocation law. The train–held-out gap makes this damage observable during training—a bridge we derive in an extended model, up to one explicitly testable condition. Building on this, we propose Gap-Adaptive Per-Objective Regularization (GAPOR), one line per objective, with provable shrinkage, rebalancing, and bounded-gain closed-loop convergence in its operative regime. Exact optimizer simulations confirm the law, and experiments across multiple LLM families with independent gold evaluation show GAPOR removes the hackable objective’s overoptimization gap while protecting true gold quality, without trading away the other objective. We also analyze the shared-channel geometry our benchmark instantiates and derive a falsifiable collapse criterion whose predicted hump shape the data confirm.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.