VWS-PFE: Modeling Vertically Weakly Ordered Structures for Pillar-Based 3D Object Detection
Abstract
Pillar-based LiDAR detectors compress the points within each vertical column into a compact feature for efficient bird's-eye-view detection. Conventional Pillar Feature Network (PFN) encoders apply shared point-wise mappings and symmetric pooling, leaving cross-point vertical relations implicit. Gravity, ground constraints, object geometry, and LiDAR sampling provide a natural organizing direction, while the observed returns remain sparse and irregularly spaced. We introduce the Vertical Weakly-Ordered Structure-Aware Pillar Feature Encoder (VWS-PFE), which realizes height-induced weak ordering through a sorted traversal and relative height-rank descriptors. Context extracted from these features is restored to each source point and reused for feature recalibration before the shared PFN projection and aggregation after projection. This connects how points are represented with how they contribute to the pillar feature, without height bins or auxiliary intra-pillar grids. With downstream detectors fixed, VWS-PFE improves mean moderate 3D AP on KITTI from 58.66 to 61.86. On nuScenes, it improves PointPillars by 4.27 NDS and 5.22 mAP, and CenterPoint-Pillar by 0.90 NDS and 1.40 mAP. KITTI ablations show benefits from contextual encoding and a further 1.43-point gain from ascending-height over random traversal. Conv1D, Transformer, and Mamba provide different accuracy–latency trade-offs within the same encoder design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.