Beyond Unidirectional Prior Injection: Bidirectional Semantic Alignment for RGB-T Semantic Segmentation
Abstract
RGB-T semantic segmentation benefits from the complementary information of visible and thermal images, while pre-trained large-scale vision models (LVMs) provide rich semantic priors for scene understanding. However, transferring RGB-pretrained semantics to RGB-T representations is nontrivial because the pre-trained semantic space is optimized for visible-image distributions and may be mismatched with heterogeneous multimodal features. To address this issue, we propose a Bidirectional Semantic Alignment Network (BiSANet), which enables target-aware interaction between RGB-T representations and pre-trained semantic priors. A Hierarchical Dual-Modal Encoder first constructs stable and complementary RGB-T features. Bidirectional Semantic Adapters then inject target-domain information into the frozen semantic pathway and transfer the adapted priors back to the multimodal branch for semantic correction. Finally, a Semantic Enhancement Decoder integrates multi-level features for fine-grained prediction. BiSANet achieves state-of-the-art performance on MFNet, PST900, and FMB, demonstrating effective semantic-prior adaptation for RGB-T segmentation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.