LMF-3DDet: Efficient LiDAR–Camera Detection via Response-Guided Sampling and Linear Fusion
Abstract
LiDAR–camera fusion enhances 3D object detection by combining geometric and semantic information, but redundant inputs and costly cross-modal interactions hinder real-time deployment. Aggressive input reduction can also discard informative observations, making it difficult to improve efficiency while maintaining detection accuracy. To address this challenge, we propose LMF-3DDet, a lightweight framework that jointly addresses input selection, feature fusion, and knowledge transfer. First, response-guided sampling uses image detection heatmaps and a geometric dilation margin to select informative LiDAR observations while retaining nearby geometric context. Second, kernelized linear cross-attention fuses the selected point tokens with image features without constructing a dense pairwise attention matrix, reducing the cost of cross-modal interaction. Third, uncertainty-aware distillation guides the lightweight student using a full-input teacher, emphasizing difficult examples without adding teacher computation at inference. Experiments on nuScenes and Waymo demonstrate a favorable accuracy–efficiency trade-off with real-time throughput, while ablation studies validate the benefits of the sampling, fusion, and distillation components.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.