LiDAR-Anchored Splat-and-Distill: Sensor-Grounded 3D-Aware Distillation for Vision Foundation Encoders
Abstract
Vision foundation encoders capture strong 2D semantics but lack the explicit 3D geometry essential for autonomous vehicle (AV) applications and embodied perception. While recent feed-forward "splat-and-distill" approaches lift camera pairs into Gaussian feature fields to distill 3D structure, they rely on inferred, scale-ambiguous geometry from frozen stereo networks, built for short-baseline, high-overlap scenes. In contrast, standard AVs are equipped with synchronized camera-and-LiDAR rigs that provide metric geometry directly. We therefore take a geometry-first approach and introduce LiDAR-Anchored Splat-and-Distill, a framework leveraging this native sensor setup to distill 3D-aware features from measured rather than inferred geometry. Adapting feed-forward distillation to AV environments raises two challenges. First, the views of a surround rig differ substantially in viewpoint, so geometry cannot be inferred from them; anchoring 3D Gaussian feature fields directly to LiDAR surfels removes any need for geometry estimation and lets distillation operate across arbitrary viewpoint changes. Second, LiDAR is sparse, confined to a narrow horizontal band, and subject to occlusion, so the lifted-and-rendered features leave most of every target view without a distillation target; a coverage-masked semantic-anchor loss preserves feature quality in those regions. An auxiliary metric-depth head supervised by the same LiDAR complements the distillation objective. We train on in-house multi-camera-and-LiDAR driving logs and evaluate zero-shot on KITTI, which uses a different sensor rig, across two backbones, DINOv2-Base and C-RADIOv4-Huge. Our method improves depth estimation and cross-view correspondence over the off-the-shelf encoders while also improving 2D semantic segmentation. Against the public SnD, FiT3D, and MEF checkpoints under the same protocol, it gives the largest depth gain and is the strongest on KITTI-pairs, a rotation-stratified driving correspondence benchmark we introduce. Our resulting encoder requires no LiDAR and adds no computational overhead at test time: measured geometry serves only as training-time scaffolding for a general-purpose 2D encoder.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.