MM-TTA: Multimodal Test-Time Adaptation with Consistency-Based Model Merging for 3D Object Detection
Abstract
Robust test-time adaptation (TTA) for 3D object detection (3DOD) plays a vital role in autonomous driving systems when facing severe domain shifts. Existing TTA methods for 3DOD mainly rely on the LiDAR modality and are vulnerable to various types of domain shifts. In this paper, we introduce a two-stage multimodal TTA method, referred to as MM-TTA, to improve the adaptation capability of 3DOD models. In the first pre-training stage, we propose the Geometric Consistency Alignment (GCA) strategy to align the 3D perception ranges of source domain multimodal data with those of the target domain, thereby preventing mismatches in parameter and feature dimensions when deploying the pretrained fusion model to the target domain, while maintaining superior performance over LiDAR-based models despite differences in camera and LiDAR configurations between domains. In the second test-time adaptation stage, we further propose the Consistency-Based Model Merging strategy (CMM) for robust transfer learning on the target domain. Instead of choosing diverse historical checkpoints, we only select the most consistent ones for model merging and use the merged model to guide the subsequent fine-tuning. This is because historical fusion models with low similarity are more likely to exhibit performance discrepancies than complementarity, which can introduce interference and degrade the merged model. To reduce the computational costs from model selection, we also propose the sign consistency score to measure the similarity between model pairs, which is computed only once without repeatedly accessing historical models. The CMM strategy achieves better performance than previous methods with fewer time steps and significantly reduces the computational costs introduced by model selection and numerous parameters of the fusion model itself. Our method is evaluated on five datasets and eight corruption types, demonstrating its effectiveness under various domain shifts in real-world scenarios (with a 15.24% improvement over the second-best method in a cross-dataset scenario).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.