X3Calib: Single Image LiDAR-Camera Calibration as Cross-Modal 3D Registration
Abstract
Extrinsic calibration between a LiDAR and a camera is a prerequisite for sensor fusion, yet learned calibrators must match 3D geometry against 2D appearance, and most are validated only on the dataset they were trained on. We reformulate calibration from a single LiDAR scan and a single image as 3D-3D registration. A monocular metric depth network lifts the image into a metric point cloud, so the LiDAR scan and the image become two point sets in one geometric domain. An overlap-aware correspondence network, trained only on KITTI, matches the two clouds, and iterated least squares with a coarse-to-fine inlier threshold turns the matches into extrinsics, with either a PnP or a 3D-3D solver. Zero-shot on nuScenes, Argoverse, and PandaSet, X3Calib outperforms LCCNet and CalibNet, neither of which even reaches the recall of simply returning the identity transform. On nuScenes, we also evaluate the recent CMRNext, which was trained on the other three datasets: X3Calib exceeds its recall by a wide margin, and CMRNext's recall again falls below that of the identity transform. A simple pipeline built on 3D registration thus generalizes better than dedicated calibration networks, even one trained on more data. The code will be made public upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.