Unity in Diversity: Low-Rank Multi-Relational Test-Time Adaptation for 3D Point Cloud Recognition
Abstract
Vision-language models have achieved strong zero-shot performance in 3D recognition, but their robustness degrades under point cloud corruptions and test-time distribution shifts. Existing point cloud test-time adaptation methods rely heavily on a single relation view, which can be unreliable under diverse corruptions. Thus, the paper highlights the global relational structure among unlabeled target samples. Specifically, we find that corrupted target relations retain compact low-rank structure, which indicates that low-rank modeling can be treated as a structural prior and a complementary view to the raw relation. Besides, when geometry is distorted, feature-space relations can be further captured by semantic relations. Raw, Low-rank, and Semantic relations together could formulate a comprehensive global multi-relational structure. Based on these observations, we propose LoReTTA, a training-free multi-relational target-set test-time adaptation framework that exploits target-set structure without updating the pretrained backbone. LoReTTA constructs multiple relation views from frozen target features and zero-shot predictions, models target dependencies through self-representation, and jointly regularizes the resulting views with a lightweight relation-mode low-rank proximal update. The refined relations are then fused into a sparse consensus graph for residual logit propagation. Experiments across multiple 3D vision-language backbones and corrupted and clean benchmarks demonstrate improved zero-shot robustness and generalization without updating the pretrained backbone. The source code will be publicly released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.