acceptodds
Under review as a conference paper at ICLR 2027

Reliability-Aware Self-Distillation for Test-Time Adaptation of Vision-Language Models

Abstract

Test-time adaptation (TTA) improves the robustness of vision-language models under distribution shifts by exploiting unlabeled test data. However, most existing methods derive supervision from the model's current predictions, which may themselves be unreliable under distribution shift and reinforce prediction errors during adaptation. We revisit masked self-distillation as a source of test-time supervision and observe that teacher reliability can be effectively indicated by consistency under mild internal perturbations. Based on this observation, we propose Reliability-Aware Self-Distillation (RASD). Specifically, RASD constructs multiple weak internal probes and aggregates their class-wise evidence into a vulnerability-aware consensus that captures both evidence strength and perturbation sensitivity. The consensus validates full-view predictions for teacher selection and provides a soft correction target, while the selected full-view predictions supervise strongly masked student views through self-distillation. To further stabilize continuous adaptation, we introduce a batch-level marginal-entropy regularization that mitigates prediction collapse. Extensive experiments demonstrate that RASD consistently improves over recent VLM-TTA methods on common corruption benchmarks, while retaining strong generalization across natural ImageNet shifts, diverse recognition datasets, and cross-domain settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.