Efficient RGB–Event Fusion for Moving Object Detection in Autonomous Driving
Abstract
In autonomous driving, low-light conditions and motion blur degrade RGB appearance representations, thereby impairing detection accuracy. RGB–Event fusion introduces event-based motion information to complement RGB appearance information when RGB imaging is degraded. However, camera ego-motion generates substantial background event responses, which interfere with target representation during cross-modal fusion. Notably, although existing multi-window methods improve detection accuracy, they struggle to balance accuracy and efficiency, ultimately constraining the reliability and real-time performance of autonomous driving perception systems. To address these challenges, we propose a single-window RGB–Event fusion network for object detection in autonomous driving. The network comprises hierarchical cross-modal fusion and multi-window-to-single-window knowledge distillation (M2S-KD). Hierarchical cross-modal fusion enhances complementary information through bidirectional feature modulation and improves deep feature fusion through Bidirectional Attention. M2S-KD transfers detection knowledge from a multi-window teacher to a single-window student through cross-head distillation, while only the student network is retained during inference to improve efficiency. Experimental results on DSEC-MOD show that the proposed method outperforms the second-best method by 2.55 percentage points in Frame [email protected], while efficiency experiments demonstrate a 115.41% increase in inference speed compared with the multi-window model. These results provide a practical approach to balancing detection accuracy and deployment efficiency for object detection in autonomous driving.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.