How Far Can Vision Test Time Training Models Adapt at Test Time?
Abstract
Test-time training (TTT) lets a model adapt its computation to every input through an inner parameter update, and how far that update can move before adaptation stops working has not been examined. We find that vision TTT adapts over a narrow range. Modest changes in update magnitude can erase the benefit of adaptation, and the preferred magnitude shifts as the input geometry changes. We show that this range is learned during training, and we test this with a minimal change to the training recipe, sampling the magnitude of the inner update on each forward pass and sharing it across all TTT blocks, so that the model learns from a continuum of adapted states while each update keeps its sample-specific direction. Range-trained models remain effective across a much broader span of update magnitudes at a small cost at the standard setting, their learned computation still depends strongly on sample-specific update directions, and the learned range persists under distribution shift, changes in input resolution and downstream fine-tuning. Vision TTT can thus learn how far to adapt at inference, which makes adaptation range a trainable property of the model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.