acceptodds
Under review as a conference paper at ICLR 2027

TopoMind: Unifying Lane Topology Reasoning and Understanding on Driving Scenes

Abstract

Lane topology reasoning provides geometric and relational constraints essential for autonomous driving. Existing methods, however, treat it primarily as a perception problem, detecting lane elements and then inferring connectivity without understanding topology-relevant scene context. We present TopoMind, to our knowledge the first unified framework for lane topology perception and topology-oriented scene understanding. TopoMind integrates lane-topology perception and topology-oriented scene understanding within a shared BEV–language architecture, where a large language model jointly processes metric BEV and text tokens and Semantic-Geometric Query Refinement couples the resulting multimodal context with geometry-anchored lane queries for topology reasoning. To supervise this capability, we assemble a topology-oriented VQA corpus comprising approximately 1.19 million question–answer pairs, covering road layout, lane markings, traffic controls, nearby road users, and coarse ego-centric directional cues, providing auxiliary language supervision over topology-relevant scene semantics. Extensive experiments on OpenLane-V2 and OmniDrive-nuScenes demonstrate that TopoMind achieves state-of-the-art performance on both lane topology perception and topology-oriented understanding, while consistently surpassing perception-only topology models. These results demonstrate the effectiveness of jointly learning topology-oriented language supervision and lane topology perception within a unified representation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.