MODE-TTA: Multi-Mode Density Modeling for Test-Time Adaptation of 3D Vision-Language Foundation Models
Abstract
3D vision–language foundation models (VLFMs) for point clouds exhibit strong zero-shot generalization capabilities, yet their performance often deteriorates under real-world distribution shifts. Test-time adaptation (TTA) seeks to address this challenge by adapting models during inference using only unlabeled target data. Recent cache-based TTA methods address distribution shifts in online settings by maintaining a memory of representative test samples and refining predictions based on feature affinities between the current sample and the cached instances. However, their effectiveness is largely dependent on the quality of pseudo-labels, which may deteriorate under substantial domain shifts. Moreover, they do not capture the underlying structure and dynamics of the incoming test distribution. To overcome these limitations, we propose MODE-TTA, a streaming density modeling framework that explicitly models class-conditional feature distributions. Each class is represented by a Gaussian Mixture Model (GMM) to capture the multi-cluster structure of 3D features, with mixture parameters updated incrementally via a hierarchical expectation–maximization procedure guided by soft zero-shot assignments. To further improve cross-modal alignment, we introduce a lightweight residual adaptation for text embeddings. Extensive experiments across diverse distribution shifts demonstrate that MODE-TTA consistently outperforms competitive baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.