Normalize the Update, Not the Weights: Efficient Loss Weighting for Multi-Objective Pretraining
Abstract
Modern deep models are often pretrained with multi-objective losses whose relative weights act as hyperparameters. Selecting these weights through random search or Bayesian optimization is computationally expensive because it requires many independent training runs. We introduce GraP, an online bilevel method that adapts pretraining loss weights to improve downstream performance through gradient alignment. A central challenge is scale ambiguity: uniformly rescaling the weights changes the magnitude of the composite update and therefore the effective learning rate. Rather than normalizing the weights themselves, GraP constrains the norm of the composite update. We show theoretically that this normalization preserves the optimal alignment direction and demonstrate its empirical benefits. To enable efficient optimization, GraP computes per-loss gradients in the shared representation space and reuses their weighted combination, reducing the number of full-backbone backward passes from K to one per iteration. Across event-sequence modeling tasks, GraP improves upon adaptive-weighting baselines and matches or improves upon carefully tuned hyperparameters, while requiring only a single training run with approximately 34% additional training-time overhead. In self-supervised computer vision, GraP remains competitive with standard tuning strategies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.