acceptodds
Under review as a conference paper at ICLR 2027

Do Cross-Modal Updates Commute? Path Dependence in Heterogeneous Multimodal Learning

Abstract

Heterogeneous multimodal learning increasingly relies on aligning signals from sensors with substantially different statistical and semantic properties. Existing work has primarily focused on what should be aligned and how strongly alignment should be imposed, whereas the effect of update order is largely obscured by the standard joint-training formulation. This paper studies whether different compositions of the same directional updates lead to the same learned solution. To isolate this question, we introduce a controlled identification protocol that fixes direction-indexed data and augmentation exposure while varying composition order. Across seven heterogeneous sensor conditions and three backbone families, reordered updates produce broad but strongly condition-dependent endpoint differences, establishing path dependence across the evaluated training settings. Complementary diagnostics examine local update compositions and endpoint sensitivity across optimizer configurations. Same-state probes detect local differences, while optimizer comparisons reveal substantial variation in endpoint sensitivity. Building on this characterization, we investigate whether path sensitivity can be controlled. Interventions at the optimizer-state, path-composition, and local-operator levels identify effective regimes, and compute-neutral antithetic block pairing reduces aggregate sensitivity by 15.5% with a small average utility change. Finally, native sequential multimodal training reproduces order-dependent endpoints across 21 fresh seeds. Fixed-scale and randomized balanced schedules yield distinct sensitivity distributions, connecting the phenomenon to schedule design in an existing alternating multimodal algorithm. Together, these results establish update order as a consequential optimization variable in sequential multimodal learning and demonstrate that schedule design can reduce its endpoint effect in the evaluated regimes while retaining matched directional exposure and the original update budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.