LiMamba: Rethinking Point Cloud Serialization in Mamba-Based Diffusion Models for LiDAR Scene Completion
Abstract
Diffusion-based methods now dominate LiDAR point cloud completion, yet every existing point-based scene-completion denoiser relies on sparse 3D convolutions whose receptive fields are inherently local, leaving long-range structural correlations largely uncaptured. Explicit global self-attention is uneconomical at outdoor scale, while state space models such as Mamba offer linear-complexity modeling but require flattening 3D voxels into a 1D sequence. Our central observation is that, while multi-directional scanning is a minor detail for object-level point clouds, at LiDAR scene scale the serialization order becomes a central design choice: the ordering governs how far the state space can propagate structure, and any single ordering is lossy. We propose LiMamba, a hybrid MinkowskiEngine–Mamba denoiser built around Multi-Scan Mamba (MSM) blocks. Each MSM block serializes voxel features along orderings derived from spinning-LiDAR acquisition physics (a radial range-density ordering and an azimuthal scan-line ordering, alongside Z-order), processes them through parallel Mamba streams, and fuses the outputs via learned gating. A Scan Dropout strategy prevents the streams from collapsing onto one during training, so that multi-scan improves accuracy rather than merely adding computation. Experiments on SemanticKITTI show that LiMamba improves over purely convolutional diffusion baselines on Chamfer Distance, JSD, and coarse-resolution occupancy IoU, while retaining linear-complexity, memory-feasible global modeling at a scale where explicit self-attention is not feasible, and ablations confirm the complementary value of multi-scan fusion over any single serialization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.