AdaMult: Adaptive Dual Optimization for Lagrangian Constrained Deep Learning
Abstract
We introduce AdaMult, a dual optimizer for Lagrangian constrained deep learning that produces feasible, well-performing solutions out of the box, at its one default hyperparameter configuration, across a wide range of tasks. This removes the costly trial-and-error tuning process of dual hyperparameters that constrained deep learning otherwise requires. AdaMult normalizes each constraint by a running estimate of its scale, so the multipliers are driven by a dimensionless signal; updates the multipliers multiplicatively, so they traverse orders of magnitude quickly; and extrapolates each multiplier along its current violation, which, under suitable step sizes and damping, makes the primal–dual dynamics locally stable at every regular local constrained minimizer. On six tasks, AdaMult at its default is feasible on every seed of every task and matches the performance of baselines that were tuned separately for each task. Held to a single configuration across tasks, every existing dual optimizer we test is infeasible on at least one of them, or feasible at a loss in performance. To our knowledge, AdaMult is the first dual optimizer that transfers reliably across tasks without tuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.