acceptodds
Under review as a conference paper at ICLR 2027

MUTE: Multimodal Uncertainty-Tempered Mutual Teaching for Lifelong Test-Time Adaptation

Abstract

Test-time adaptation (TTA) is crucial for deploying multimodal models in real-world scenarios where test distributions differ from training data. Multimodal test-time adaptation (MMTTA) further leverages complementary information across multiple modalities to improve robustness. In this work, we focus on audio-video MMTTA, where models process temporally aligned audio and video inputs to produce audio, video, and fused predictions. However, we observe that existing TTA methods suffer significant performance degradation in lifelong MMTTA settings, where models must adapt sequentially to a stream of diverse corruptions without forgetting previous knowledge. To address this challenge, we propose MUTE (Multimodal Uncertainty-Tempered Mutual Teaching), a plug-and-play framework that uses confidence-based routing with a modality-advantage criterion to select the most reliable teacher among the audio, video, and fused heads and transfer its supervision to the others. Confidence-conditioned temperature scheduling and adaptive weight modulation stabilize mutual teaching, while our analysis provides a margin-dependent per-step risk bound and a temperature-dependent teacher-target sensitivity bound. MUTE is buffer-free and composes seamlessly with existing TTA methods. Extensive experiments with twelve base TTA methods across two datasets and both audio and video corruption streams demonstrate that MUTE generally improves average accuracy, with gains of up to 42.9 percentage points for severely degraded base methods and up to 2.1 points for already stable ones. Code is available at https://anonymous.4open.science/r/MUTE-ICLR27-C2B4/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.