Can Hybrid-Space Representation Learning Advance Multimodal Fusion?
Abstract
Multimodal fusion—joint learning of representations from heterogeneous data has shown significant potential in healthcare by integrating imaging, text, and omics for comprehensive analysis. However, prior methods face two major limitations: (1) inadequate modeling of complex hierarchical structural dependencies at high computational cost, which constrains the performance–efficiency trade-off ; and (2) limited robustness to missing modalities. These limitations restrict their applicability in resource-constrained medical AI. To address these challenges, we introduce Hybrid Space Attention Refiner (HySAR), a plug-and-play cross-modal refiner that jointly exploits Euclidean and non-Euclidean manifolds to capture complex hierarchical structural dependencies for learning robust shared representations. This, in turn, improves multimodal fusion performance and missing-modality robustness, while HySAR’s channel-wise geometric refinement and complementary multi-scale information learning further reduce computational cost. Across 10 multimodal benchmarks, HySAR-integrated learners achieve ≈ 4.8-point on average improvements over baselines, up to ≈ 84.1% parameter and ≈ 88.0% FLOP reductions, and up to ≈ 8.7-point gains in missing-modality robustness, enabling efficient and robust clinical prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.