acceptodds
Under review as a conference paper at ICLR 2027

MapDLM: Lane-Aligned Block Diffusion for Lane Graph Generation

Abstract

Constructing HD maps from multi-frame LiDAR BEV representations requires predicting lane-centerline geometry and directed lane connectivity. Generic vision–language model tokenizers are not designed to represent coordinates, while long token sequences weaken the association between coordinate points and the lane geometry and topology they define. We present MapDLM, a lane-supervised multi-token generation model based on lane-aligned block diffusion. MapDLM first predicts a lane-ID list and directed lane connectivity, then reconstructs each lane as fixed-length token blocks conditioned on this structure. In addition to the standard cross-entropy loss, training combines point-level coordinate objectives, lane-level polyline constraints, and topology-chain supervision. Anchor-Point Gaussian Loss (APG Loss) tolerates errors between nearby coordinate bins in lane interiors while retaining exact targets at lane endpoints. Semantic-Adaptive Noise Scheduling and structure-aware commitment expose the model to the highly masked states encountered during iterative denoising. By jointly predicting multiple lanes within each diffusion block, MapDLM reduces serial decision depth and error propagation. We evaluate MapDLM on ArgoverseĀ 2 and nuScenes under the OpenLane-V2 protocol, compare it with corresponding autoregressive baselines, and conduct controlled ablations of lane alignment, multi-level supervision, and state-aligned denoising. The results show effective geometry and topology reconstruction on both benchmarks and support lane-aligned diffusion as a promising approach to structured map generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.