TokenI2P: Learning Pose-Aware Tokens for Correspondence-free Image-to-Point Cloud Registration
Abstract
Image-to-point cloud registration aims to estimate the 6-DoF camera pose between a 2D image and a 3D point cloud, which remains highly challenging due to the substantial modality gap. Existing methods, whether correspondence-based or correspondence-free, largely rely on multi-stage pipelines with explicit matching or intermediate proxy predictions. Consequently, early-stage errors unavoidably propagate to subsequent pose solvers, limiting robustness while retaining high pipeline complexity. In this paper, we propose TokenI2P, a novel single-stage, correspondence-free framework that reformulates 2D-3D registration as pose-aware cross-modal compression. Specifically, TokenI2P introduces learnable pose tokens that dynamically interact with heterogeneous 2D image and 3D point cloud features via cross-modal attention, progressively compressing pose-discriminative evidence into a compact representation for direct 6-DoF pose regression. By doing so, it completely bypasses explicit matching, intermediate geometric proxies, and RANSAC-like optimizations in a unified feed-forward pass. To guide this compression, we maximize cross-modal mutual information on the 2D/3D features (attended by pose tokens) at both frame and point levels. Frame-level mutual information enforces global overlap awareness, while point-level mutual information promotes fine-grained geometric discrimination between corresponding point-pixel pairs, together compelling the network to filter out modality-specific noise. Thus, pose-token aggregation governs how cross-modal observations are compressed, while mutual-information learning dictates what spatial evidence is preserved. Extensive experiments on the KITTI and nuScenes benchmarks demonstrate that TokenI2P achieves competitive registration accuracy with significantly lower pipeline complexity and superior inference speed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.