Dimension-wise Distributionally Robust Optimization for Multimodal Learning via Adaptive Aggregation
Abstract
Multimodal data inherently carries multiple semantic dimensions, such as imaging modality, anatomical region, and task type in medical data. Each dimension induces subpopulation shifts that may vary independently. Group Distributionally Robust Optimization (Group DRO) provides a principled framework for addressing such shifts by minimizing worst-case losses across predefined groups. However, when multimodal data carries multiple grouping dimensions, the standard practice of combining them via their Cartesian product leads to an explosive growth in the number of groups (for instance, combining 8 modalities, 5 anatomical regions, and 5 task types yields 200 groups), leaving each group with extremely sparse samples and making robust optimization unstable. To address this, we propose Dimension-wise DRO, which formulates robustness along each semantic dimension independently and adaptively aggregates dimension-wise risks into a unified objective. Our method consists of three components: (i) selecting non-redundant dimensions by pruning highly correlated axes; (ii) defining a smooth robust risk for each selected dimension via KL regularization; and (iii) aggregating dimension-wise risks with a Hedge-style adaptive weighting scheme that prioritizes harder dimensions. We theoretically prove that the proposed aggregated objective bounds the worst-case group risk across all dimensions, with a convergence rate of , matching standard SGD up to a logarithmic factor. Experiments on three multimodal benchmarks—OmniMedVQA, VQAv2, and BDD100K—demonstrate that our method consistently improves worst-group performance under both shifted and unshifted settings without sacrificing average accuracy, outperforming both annotation-dependent and annotation-free baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.