MapDLM: Lane-Aligned Block Diffusion for Lane Graph Generation
Abstract
Constructing HD maps from multi-frame LiDAR BEV representations requires predicting lane-centerline geometry and directed lane connectivity. Generic visionālanguage model tokenizers are not designed to represent coordinates, while long token sequences weaken the association between coordinate points and the lane geometry and topology they define. We present MapDLM, a lane-supervised multi-token generation model based on lane-aligned block diffusion. MapDLM first predicts a lane-ID list and directed lane connectivity, then reconstructs each lane as fixed-length token blocks conditioned on this structure. In addition to the standard cross-entropy loss, training combines point-level coordinate objectives, lane-level polyline constraints, and topology-chain supervision. Anchor-Point Gaussian Loss (APG Loss) tolerates errors between nearby coordinate bins in lane interiors while retaining exact targets at lane endpoints. Semantic-Adaptive Noise Scheduling and structure-aware commitment expose the model to the highly masked states encountered during iterative denoising. By jointly predicting multiple lanes within each diffusion block, MapDLM reduces serial decision depth and error propagation. We evaluate MapDLM on ArgoverseĀ 2 and nuScenes under the OpenLane-V2 protocol, compare it with corresponding autoregressive baselines, and conduct controlled ablations of lane alignment, multi-level supervision, and state-aligned denoising. The results show effective geometry and topology reconstruction on both benchmarks and support lane-aligned diffusion as a promising approach to structured map generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.