X-MV3D: Benchmarking Cross-Dataset Generalization in Multi-View 3D Object Detection
Abstract
Despite substantial progress in multi-view 3D object detection, existing detectors are still developed and evaluated largely within individual datasets, leaving their ability to transfer to newly collected datasets unclear. We introduce X-MV3D, a unified benchmark for cross-dataset scene-level 3D detection from multi-view captures. X-MV3D converts ARKitScenes, ScanNet, HyperSim, and 3RScan into a common gravity-aligned coordinate frame, with standardized scene records, category mappings, and 3D box annotations, enabling the four datasets to be trained and evaluated under the same protocol. We define three evaluation settings: single-source transfer, leave-one-dataset-out (LODO) transfer, and target-included mixed training, with LODO serving as the primary target-unseen setting. Across representative detector families, we find that current multi-view 3D detectors are strongly dataset-conditioned. Multi-source training improves over single-source transfer, but LODO remains far below target-included training, showing that multi-source training alone is insufficient for reliable new-dataset transfer. Failure decomposition further shows the transfer gap is dominated by degraded proposal coverage and box localization, suggesting that acquisition shifts may affect the camera-dependent feature lifting and multi-view aggregation. Motivated by this analysis, we propose CamGeoGate, a lightweight detector-agnostic baseline that injects explicit camera geometry into image features and trains detectors with stochastic view gating. CamGeoGate consistently improves LODO performance across three detector families, showing that camera-aware and view-robust training is a promising direction for reusable multi-view 3D detectors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.