acceptodds
Under review as a conference paper at ICLR 2027

Decoupling to Fuse: Learning Shared and Modality-Specific Representations for Multi-Sensor 3D Object Detection

Abstract

Fusing information from cameras, LiDAR, and 4D radar is an important approach to improving 3D object detection in autonomous driving. However, multimodal features contain a mixture of shared information, modality-specific cues, and background responses. A key challenge is to exploit cross-modal consistency while preserving useful differences and guiding their fusion according to object relevance. We introduce ObjDec, a multi-sensor fusion architecture that learns shared and modality-specific representations and uses object-related information to guide their fusion. Specifically, ObjDec applies constraints to aligned local representations within foreground regions, encouraging cross-modal consistency in shared features while preserving differences among modality-specific features. Based on these representations, foreground gating and modality control regulate each modality’s feature contribution, while object context further modulates cross-modal attention queries and fused outputs. Experiments on K-Radar dataset demonstrate that ObjDec achieves state-of-the-art performance, improves 8.04% @IoU=0.3 (88.36%)and 7.77% @IoU=0.5 (88.10%)at a confidence threshold of 0.3, with only approximately 1.7% more parameters and fast inference. Component ablations and analyses of learned representations and spatial gating further support the effectiveness of the proposed design. Our code will be made publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.