Rethinking Unaligned RGB-T Dense Prediction as Cross-Modal Information Transport
Abstract
RGB-T dense prediction remains challenging because visible and thermal observations are rarely perfectly aligned in practice. When the two modalities are spatially inconsistent, the thermal response at a given coordinate may originate from a different scene region, and reweighting such co-located responses can only attenuate the mismatch rather than recover the complementary information that lies elsewhere. We therefore propose unaligned cross-modal information transport (UCMIT), which reformulates RGB-T interaction as a structured cross-modal information transport process. Within UCMIT, local cross-modal information transport (LCIT) constructs a semantic-geometric transport plan through entropy-regularized optimal transport, retrieving complementary thermal evidence from plausible nearby locations and reconstructing it in the visible coordinate system without explicit global registration. Source-side marginal constraints couple evidence allocation across RGB targets and prevent unrestricted reuse of dominant thermal responses, while a fixed window-level null mass is redistributed across targets according to their relative matching preference. Reliability-controlled evidence fusion (REF) then arbitrates between the co-located and transported responses according to the local transport state and regulates residual injection according to evidence reliability and path-aware certainty. LCIT thus determines where complementary evidence is obtained, while REF determines which evidence is retained and how strongly it modifies the visible representation. Extensive experiments on RGB-T salient object detection and semantic segmentation demonstrate strong overall performance across unaligned benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.