When and Why Can Test-Time Adaptation Succeed?
Abstract
Test-time adaptation (TTA) modifies deployed models using only unlabeled target data, yet a fundamental theoretical question remains open: when can unlabeled observations determine a better target predictor, and why can entropy-based up- dates correct errors rather than merely sharpen them? We identify recoverable semantic structure as a common principle underlying both questions. First, we characterize conditions under which target semantics are recoverable from unla- beled data: semantic components must be recoverable and oriented by the source model, with within-semantic support dominating cross-semantic leakage. Under these conditions, we establish pointwise component-resolution learnability in the probably approximately correct (PAC) sense and derive finite-sample guarantees for exact component-level semantic identification. We then ask how entropy-based TTA algorithms exploit this information. For finite-width two-layer networks with a fixed output head, we show that entropy on an isolated prediction carries no cor- rectness signal; correction instead arises from parameter sharing, which propagates source-oriented evidence through a parameter-gradient coupling kernel. Under a positive support-over-leakage condition controlling finite-width dynamics, we prove early-stage target-risk reduction. Guided by the same principle, we finally introduce PRojected Invariant Spectral Matching (PRISM), an optimization-free method that explicitly preserves shared task-relevant evidence while suppressing view-specific nuisance in the classifier-visible subspace. Across five distribution- shift benchmarks, PRISM consistently outperforms strong backpropagation-based and backpropagation-free TTA baselines. Code is available here.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.