acceptodds
Under review as a conference paper at ICLR 2027

SACFusion: From Semantic Compatibility to Conditional Complementarity in RGB-Event Object Detection

Abstract

RGB-Event object detection leverages complementary information to improve robust perception under challenging and highly dynamic conditions, particularly for autonomous driving. However, substantial differences between RGB and Event modalities in semantic representation, observation density, and information reliability often lead to modality bias during fusion. To address this issue, we propose SACFusion, a two stage RGB-Event object detection framework that progresses from semantic compatibility to conditional complementarity. In Stage I, a Temporal Polarity Aggregator explicitly models the temporal, polarity, and local spatial structures of Event data, while Scale Aware Local Semantic Alignment and Cross Path Semantic Preservation adapt Event features to the pretrained RGB semantic space without reconstructing RGB appearance. In Stage II, an RGB dominant asymmetric fusion mechanism uses a spatially varying Event Retrieval Prior to guide RGB queries toward more informative Event regions at corresponding feature scales. This design preserves stable RGB representations while reducing interference from sparse or unreliable Event observations. By separating representation adaptation from complementary interaction, SACFusion enables Event information to be exploited more selectively according to local scene conditions. Experiments on DSEC-Det and PEOD demonstrate consistent improvements under diverse and challenging conditions, achieving relative mAP gains of 11.1% and 16.5%, respectively, over the strongest competing methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.