acceptodds
Under review as a conference paper at ICLR 2027

Continuous - Time Voxel free Image - Event Semantic Segmentation.

Abstract

Semantic segmentation for autonomous driving and robotics remains challenging under fast motion, low illumination, and high dynamic range, where conventional RGB sensing becomes unreliable. Event cameras provide complementary measurements with high temporal resolution, yet most existing RGB-event segmentation methods first quantize asynchronous events into voxel grids, discarding fine-grained temporal structure and potentially introducing cross-modal misalignment. We present a voxel-free RGB-event segmentation framework that preserves event timing throughout the representation and fusion process. At its core is CT-EvSSM, a continuous-time event encoder that operates directly on raw events using per-event temporal embeddings and a selective state-space model whose discretization is conditioned on the true inter-event intervals. To integrate event features with RGB representations, we adapt a pretrained DINOv2-L/14 backbone into a multi-scale feature pyramid and introduce a dual-branch fusion mechanism that combines geometric cross-modal alignment with frequency-domain feature fusion and temporal state-space modeling. This design enables event information to be aligned and integrated with RGB features while retaining its asynchronous temporal structure. Extensive experiments on two widely used datasets, DSEC and DDD17, demonstrate the effectiveness of the proposed method compared to several state-of-the-art approaches.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.