acceptodds
Under review as a conference paper at ICLR 2027

Minimax Update Balancing for Unified Multimodal Pretraining

Abstract

Unified multimodal pretraining optimizes image understanding (I2T) and generation (T2I) through one shared backbone and one shared optimizer, and balanced task sampling does not balance the optimizer, since under SOAP the realized update serves the two directions unequally. The optimizer averages gradients over accumulation segments rather than balancing tasks directly. As a result, the averaged preconditioned update can be dominated by the segment with the larger norm under the preconditioner's metric. To address this imbalance, we introduce ML-FOP-SOAP, which operates within each gradient-accumulation window. It hierarchically combines consecutive segment averages in a bottom-up manner. At each fold, it computes the difference between the two segment averages, extracts the component orthogonal to the SOAP update, and adds this component back with a coefficient determined by minimax update balancing. This coefficient minimizes the worst-case local surrogate cost across the two segments, preventing either segment from dominating the update. The coefficient has a bounded closed-form solution. It is zero when the two segments are already balanced, equalizes their local surrogate costs when feasible, and otherwise falls back to the standard SOAP update. When the costs can be equalized, the correction follows the preconditioned multiple-gradient descent direction. Importantly, the preconditioner is estimated from the uncorrected mean gradient, keeping the balancing correction separate from preconditioner estimation. On Janus-400M, ML-FOP-SOAP achieves the lowest T2I loss among all tested optimizers and balancing methods while also reducing I2T loss relative to SOAP. It improves the entire I2T–T2I loss frontier across sampling ratios. Its aggregate loss improvement over SOAP is comparable to SOAP's improvement over AdamW, and it matches Muon's performance. Ablation studies show that each design choice matters. Using a fixed coefficient or Euclidean inner products performs worse than SOAP, while applying a single correction over the entire accumulation window captures only part of the improvement. Explicitly contrasting I2T and T2I at every fold degrades performance in both directions, supporting our segment-based approach. Across all three tested architectures, ML-FOP-SOAP achieves the highest mean proportion of updates that benefit both directions, indicating more balanced optimization. Its T2I improvement also increases when scaling from 400M to 1B parameters, where it achieves a 12% improvement in FID-1K.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.