Mind the Gap Between Train and Test Grids: Resolution Collapse in Voxel-Indexed Point Networks Is Mostly a Units Problem
Abstract
For point cloud semantic segmentation, the leading backbones voxelize the scene and read it along a space-filling curve, which keeps their cost low on large scenes. Evaluated on a grid coarser than the one they trained on, they lose around 50 mIoU points on standard indoor and outdoor benchmarks. The usual explanation is that a coarser grid destroys geometry. In this work, we show that most of that collapse reverses without restoring any geometry. Sparse convolution, pooling, serialization and rotary position encoding all read the integer index , so changing the voxel size changes the unit of every operator, and nothing in the network signals it. We propose relabelling: it recomputes that index from the physical coordinate at the training grid, without retraining. On six architecture-dataset pairs, it recovers 70 to 90% of the collapse on grids three times coarser than the training one. On dense indoor scenes, it cuts peak memory by 3.4 to 4.4× and inference time by up to 46%. Most of what a coarser grid appears to cost is a unit mismatch: even on grids finer than the training one, which lose no geometry, the network still collapses, and our method recovers all of it. This new training-free relabelling step offers a generic speedup for any voxel-indexed point network.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.