Intervention-Dependent Selection Credit in Stateful Test-Time Adaptation
Abstract
Selective test-time adaptation (TTA) methods choose which test samples drive each update. Their gains are often credited to reliable samples, but selection also changes sample counts and loss weights, while each update changes the model that scores later samples. Standard ablations conflate these effects. We measure the value of sample identity by recording native per-stage counts and rerunning each stream with those counts, choosing samples by each run's own native ranking or at random. Crossing these choices at the two stages of DeYO, EATA, and SAR yields a factorial that splits identity value between stages and measures their interaction. Native selection has positive identity value in 164 of 180 ImageNet-C cells. Yet stage credit changes with the comparison: on a five-corruption panel with 10 seeds, first-stage dominance is resolved in 31 of its 35 cells under the original comparison; matching mean loss weights leaves 19 of 40 cells first-stage dominant, with 8 resolved reversals (9 with joint resampling), and removes most of the ViT-B/16 identity value. In 8 of 12 candidate SAR cells, 20-seed intervals show that stages helping individually can hurt together. Waterbirds further separates identity-driven worst-group gains from harms that weaken under uniform loss weights. Selection credit is therefore intervention-dependent: it describes a stage's marginal effect within a specified stateful comparison, rather than an intrinsic property of its criterion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.