acceptodds
Under review as a conference paper at ICLR 2027

GeoAct: Queryable Geometric Representations for Autonomous Driving

Abstract

Driving stacks need scene geometry that is metric at inference time and that downstream modules can read without re-running a reconstruction. We present GeoAct, a calibration-aware encoder that turns multi-camera image histories into one shared geometric memory. The memory is read in two ways: densely, by a DPT-style head predicting per-pixel metric depth and point maps, and sparsely, by a point query asking where the surface seen at a pixel and time lies at a requested time. The encoder has about 100M parameters and is trained from projected LiDAR and monocular pseudo-labels, without distilling a large 3D model. On the NAVSIM/OpenScene test split it reaches 0.1093 AbsRel with no test-time scale alignment, more accurate than GeoX even with a per-scene oracle scale. Behind a frozen encoder, a trained adapter and DrivoR-based planner reach 91.78 PDMS, above a frozen GeoX memory under the same planner budget (90.83). Controlled continuations from one checkpoint reveal a trade-off: dropping dense supervision improves observed-motion prediction and degrades current-frame geometry. Under our frozen-transfer protocol, these two variants show no detectable planning difference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.