Learning a Structured Latent Space for Vectorized HD Maps
Abstract
Vectorized HD maps provide structured geometric and semantic descriptions essential for autonomous driving. As autonomous systems advance toward generative world models and simulation platforms, maps are increasingly required to function as learnable, generative interfaces rather than static offline assets. However, they naturally take the form of variable-length, heterogeneous, and unordered sets of geometric primitives, such as lane polylines and road polygons. Integrating such irregular vector collections into modern continuous generative models presents a fundamental structural challenge. Existing approaches typically struggle with a trade-off: imposing heuristic one-dimensional sorting on primitives induces semantic drift that collapses outputs into dataset-average straight corridors, whereas uncoordinated set diffusion fails to organize global scene layouts, yielding trivial, mutually parallel paths. To resolve this dilemma, we propose MapVAE, a hierarchical autoencoding framework that decouples local geometric modeling from global scene aggregation into a compact continuous latent space with canonical slot identities. In the first stage, a unified Vector Autoencoder projects heterogeneous primitives onto a shared hyperspherical manifold using canonical local coordinates, faithfully preserving fine-grained curvatures without category-specific backbones. In the second stage, a Map Autoencoder aggregates these element-level features into a fixed-length bottleneck via learnable queries, establishing structured slot representations that anchor global layout and directional diversity for downstream continuous generation. Crucially, distinct from existing works that decode raw geometries, we formulate the Map Autoencoder as a DETR-style set reconstruction directly in feature space, which unifies heterogeneous primitive recovery into a shared metric and improves optimization stability via unit-hypersphere constraints, a discriminative triplet loss, and spherical geodesic denoising. Extensive experiments demonstrate that MapVAE achieves superior reconstruction fidelity and enables our diffusion model, MapLDM, to synthesize diverse, geometrically precise HD maps across complex urban scenes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.