acceptodds
Under review as a conference paper at ICLR 2027

ACME: Online Adaptation via Cross-Moment Estimation for Vision-Language Models

Abstract

Vision-language models (VLMs) provide strong zero-shot predictions, but their performance can deteriorate under test-time distribution shift. Test-time adaptation (TTA) aims to mitigate this degradation using unlabeled test samples during inference. Existing cache-based TTA methods retain confidence-selected examples or hard pseudo-labels, which can discard useful target evidence and propagate unreliable predictions. Distributional methods instead estimate target-domain feature distributions, but often require costly covariance state or task-sensitive hyperparameters whose improper configuration can cause severe performance degradation. To address these issues, we introduce Adaptive Cross-Moment Estimation (ACME), a tuning-free and backpropagation-free online adapter that uses the complete frozen posterior of every test sample. ACME maintains exact streaming feature and posterior statistics together with their cross-covariance, from which it derives class-specific directions for a bounded residual correction to the frozen logits. ACME requires no source data, stored test samples, or full class-conditional covariance matrices. Extensive experiments across distribution-shift and cross-dataset benchmarks demonstrate that ACME achieves strong adaptation performance with a negligible additional GPU memory overhead of 3 MiB and the highest inference throughput among the evaluated TTA methods, retaining 88% of frozen CLIP's throughput on ImageNet-1K.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.