Geometry Before Fusion: Event-First State-Space Modeling for Event-Frame Stereo
Abstract
Event cameras remain informative under fast motion and extreme illumination but are sparse, whereas frame cameras provide dense appearance that degrades under the same conditions. Existing event–frame stereo methods fuse the two modalities before or during matching, so unreliable appearance can bias correspondence before any geometric reference exists. We propose to establish geometry before multimodal fusion: disparity is first estimated solely from stereo events with a state-space architecture, then refined by a single frame weighted by a reliability measure derived from event stereo consistency. On DSEC, our method is the most accurate among methods using at most one frame camera and approaches symmetric methods using two, while its event-only variant outperforms prior event-only methods. It also generalizes zero-shot to MVSEC and M3ED. Ablations show that fusion timing is critical: introducing the frame before matching degrades accuracy below the event-only model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.