Robust Test-Time Text–3D Scene Retrieval: Benchmark and Reliability-Routing
Abstract
Deployed Text–3D Scene Retrieval (T3SR) systems receive unlabeled queries that deviate from the training distribution: typos and paraphrases on the text side, and occlusion, sensor noise, and partial scans on the 3D side. Test-time adaptation (TTA) can help address these distribution shifts, but its effectiveness depends on the reliability of self-generated cross-modal correspondences. Two coupled challenges arise: gallery-level matching biases can undermine correspondence reliability despite high prediction confidence, while shared query-encoder updates can induce representation drift in uncertain queries even when they are excluded from the adaptation objective. To systematically study these challenges, we introduce M3SP, a benchmark comprising 14 multi-level 3D perturbation types, each instantiated at five severity levels across four datasets. We further propose **STAR-3D** (**S**tructure-guided **T**est-time **A**daptation for Text–3D Scene **R**etrieval). STAR-3D first calibrates query–gallery matching scores through history-aware Sinkhorn normalization, then assesses provisional correspondences by comparing query and candidate representations in a centered residual space. Structural consistency and a memory-adaptive confidence gate jointly determine which queries drive confidence sharpening, while the remaining queries are regularized toward their frozen-source representations. Extensive experiments across datasets demonstrate that STAR-3D outperforms competitive TTA baselines overall under query shift and query–gallery shift. Code is available at [https://anonymous.4open.science/r/STAR-3D-C486](https://anonymous.4open.science/r/STAR-3D-C486).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.