acceptodds
Under review as a conference paper at ICLR 2027

Dual-diffusion Spatial Interaction Representation Generalization for 3D Semantic Segmentation

Abstract

3D semantic segmentation in unknown environments is crucial. Current methods utilize cross-modal learning and introduce 2D semantic priors to improve domain-generalized 3D segmentation. However, the heterogeneous representations of different modalities hinder cross-modal interaction, limiting the adaptation of 2D semantics to 3D space. In this paper, Dual-diffusion Spatial Interaction Representation (DSIR) is proposed to explore how cross-modal diffusion representations enhance 3D spatial semantics to improve the generalization of 3D semantic segmentation. Specifically, Spatial Representation Alignment (SRA) is first proposed to align the cross-modal diffusion representations, which designs manifold alignment of attributes and spatial coordinates to promote discriminativeness and spatial consistency before modalities interact. Secondly, Bidirectional Interaction Representation (BIR) models dual-diffusion interactions, which designs spatial joint adaptive representations to enable 2D and 3D bidirectional interactions. Finally, Spatial Semantic Modulation (SSM) is proposed to enhance 3D spatial semantics through effective interaction representations, which modulates local semantics in 3D space while suppressing spatially uncorrelated interaction noise, thereby improving the generalization of 3D semantic segmentation. Extensive experiments demonstrate that the proposed method has state-of-the-art performance. The code will be available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.