acceptodds
Under review as a conference paper at ICLR 2027

Can Hybrid-Space Representation Learning Advance Multimodal Fusion?

Abstract

Multimodal fusion—joint learning of representations from heterogeneous data has shown significant potential in healthcare by integrating imaging, text, and omics for comprehensive analysis. However, prior methods face two major limitations: (1) inadequate modeling of complex hierarchical structural dependencies at high computational cost, which constrains the performance–efficiency trade-off ; and (2) limited robustness to missing modalities. These limitations restrict their applicability in resource-constrained medical AI. To address these challenges, we introduce Hybrid Space Attention Refiner (HySAR), a plug-and-play cross-modal refiner that jointly exploits Euclidean and non-Euclidean manifolds to capture complex hierarchical structural dependencies for learning robust shared representations. This, in turn, improves multimodal fusion performance and missing-modality robustness, while HySAR’s channel-wise geometric refinement and complementary multi-scale information learning further reduce computational cost. Across 10 multimodal benchmarks, HySAR-integrated learners achieve ≈ 4.8-point on average improvements over baselines, up to ≈ 84.1% parameter and ≈ 88.0% FLOP reductions, and up to ≈ 8.7-point gains in missing-modality robustness, enabling efficient and robust clinical prediction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.