ROWA: A Low-Complexity Adaptive Weight Framework for Multi-Task Learning
Abstract
Gradient manipulation methods in Multi-Task Learning (MTL) are effective for balancing competing objectives, yet they often suffer from severe computational bottlenecks. Recent -aware utility optimization approaches unify various task-level fairness criteria, but require expensive task-specific backward passes at each training step, where denotes the number of tasks. To overcome this scalability barrier, we propose ROWA (GeneRic Online Weight Adaptation), which is derived from an online optimization formulation incorporating the loss decrease and -aware targets. ROWA is implementation-friendly governed by a single mirror descent step at each iteration, and completely obviates the need for any task-specific gradient computation during task weighting. Remarkably, ROWA can be generalized to arbitrary -configurations and readily incorporate the Weighted-Chebyshev strategy to enable exploration of user-preferred regions on the Pareto front. We establish an iteration complexity for the convergence of ROWA to Pareto stationarity in the nonconvex setting. Extensive experiments on four benchmarks with two to 40 tasks show that ROWA outperforms a wide range of strong baselines at negligible computational overhead. Ablation studies further confirm its flexibility across fairness criteria and Pareto-front preferences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.