Decoupling to Fuse: Learning Shared and Modality-Specific Representations for Multi-Sensor 3D Object Detection
Abstract
Fusing information from cameras, LiDAR, and 4D radar is an important approach to improving 3D object detection in autonomous driving. However, multimodal features contain a mixture of shared information, modality-specific cues, and background responses. A key challenge is to exploit cross-modal consistency while preserving useful differences and guiding their fusion according to object relevance. We introduce ObjDec, a multi-sensor fusion architecture that learns shared and modality-specific representations and uses object-related information to guide their fusion. Specifically, ObjDec applies constraints to aligned local representations within foreground regions, encouraging cross-modal consistency in shared features while preserving differences among modality-specific features. Based on these representations, foreground gating and modality control regulate each modality’s feature contribution, while object context further modulates cross-modal attention queries and fused outputs. Experiments on K-Radar dataset demonstrate that ObjDec achieves state-of-the-art performance, improves 8.04% @IoU=0.3 (88.36%)and 7.77% @IoU=0.5 (88.10%)at a confidence threshold of 0.3, with only approximately 1.7% more parameters and fast inference. Component ablations and analyses of learned representations and spatial gating further support the effectiveness of the proposed design. Our code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.