FreqLift: BEV Fusion and Cross-Frequency State Space Modeling for LiDAR-Camera 3D Object Detection
Abstract
In recent years, significant progress has been made in BEV-based 3D object detection. However, processes such as voxelization, viewpoint transformation, and downsampling can still compromise local structural details. Camera features, when projected into the BEV space, are affected by depth estimation and projection errors, and direct fusion may introduce these errors into the LiDAR-camera BEV representation. To address these issues, this paper proposes FreqLift, a LiDAR-camera 3D detection framework that combines BEV fusion with cross-frequency state-space modeling. FreqLift first refines the camera BEV using CameraBEVRefiner and fuses it with the LiDAR BEV. Subsequently, WaveLift uses multi-level wavelet decomposition and Wavelet Cross-Gate to model high- and low-frequency interactions; Freq-Mamba then uses low-frequency structural differences and BEV position coordinates to modulate the state-space model, thereby enhancing the structural modeling capabilities of the fused BEV. On the nuScenes validation set, FreqLift achieves 71.2% mAP and 73.1% NDS, whereas its LiDAR-only variant, FreqLift-L, achieves 67.1% mAP and 71.0% NDS. An ablation shows that this +4.1 mAP gain comes from direct camera fusion (+1.4 mAP) and CameraBEVRefiner (+2.7 mAP).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.