acceptodds
Under review as a conference paper at ICLR 2027

GeoX-2: High-Throughput Driving Geometry through Geometry-Preserving Compression

Abstract

Feed-forward visual geometry models jointly recover camera pose, depth, and dense 3D structure, but their computational cost limits efficient scaling to synchronized driving views. We present , which compresses the complete inference graph for efficient driving geometry. Structurally, an Efficient Hybrid Encoder reduces the token count of each view, while a lightweight dense head supports efficient geometry inference. To compress the decoder, GeoX-2 treats each Frame–Global block pair as a removable unit and applies Geometry-Derived Structural Slimming. The Pair Importance Score combines LayerScale-based structural saliency with depth-intervention sensitivity to retain important pairs from the full decoder. Prior-free geometry distillation then transfers geometry from a prior-conditioned teacher to the selected compact RGB-only student, removing the conditional adapters and learned point head while recovering 3D points from predicted depth and camera poses. Extensive experiments across multiple driving benchmarks demonstrate strong geometry accuracy. GeoX-2 reaches 1003.1 FPS offline sequence-batched throughput on KITTI and provides leading throughput and peak-memory efficiency across multi-view inference settings with varying view counts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.