acceptodds
Under review as a conference paper at ICLR 2027

Uncertainty-Aware Multimodal Fusion for Robust Urban Representation

Abstract

Urban representation learning increasingly relies on multiple sensing modalities to capture the diverse functional, behavioral, and physical characteristics of geographic regions. Existing methods improve multimodal representations through cross-view alignment or adaptive fusion, but largely treat each observed modality as deterministic evidence whose reliability is implicit in its features. This assumption is fragile in real urban sensing, where data quality and coverage can vary substantially across both modalities and regions. As a result, unreliable observations may distort cross-view semantic learning and propagate into the final representation. We introduce URGE, an uncertainty-aware framework built on a simple principle: a modality should be represented by both what it suggests about a region and how certain that suggestion is. URGE learns probabilistic view-specific semantics through Pairwise Gaussian Agreement, which evaluates cross-view compatibility relative to uncertainty without requiring different modalities to share the same uncertainty level. It then carries uncertainty into fine-grained token-level fusion, allowing reliability to condition how semantic evidence interacts across modalities. Finally, we introduce Relative Fusion Robustness, which uses controlled evidence degradation and counterfactual fusion to encourage the model to make productive use of uncertainty rather than merely predict it. Across real-world urban benchmarks covering multiple cities and downstream tasks, URGE yields more effective region representations and remains more robust when multimodal evidence is degraded.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.