The Features We Make Along the Way: Intermediate Representations for Training-Free Continual Learning
Abstract
Intermediate representations can improve frozen-backbone classification, but combining depths, selecting a better layer, and changing the classifier can each contribute. We examine these choices across eleven image datasets and four frozen ViTs, comparing layer selection, normalization, and classifier heads on the same features. On five datasets covering digits, traffic signs, and overhead imagery, normalized all-layer features outperform the best validation-selected single block on 18 of 20 backbone-dataset pairs with KLDA and 19 with RanPAC. Equally wide random nonlinear views of the final block do not recover these gains. Direct ridge regression on raw all-layer features performs best on these datasets, reaching 91.82% accuracy, compared with 88.94% for normalized RanPAC and 88.97% for normalized KLDA. Per-block normalization improves the random-feature heads but leaves ridge accuracy essentially unchanged. With hyperparameters selected without future-task data, all-layer inputs add approximately four percentage points of average incremental accuracy for both random-feature heads on the same five datasets, although restricted baseline tuning likely inflates RanPAC's gain. The dataset grouping was exploratory. A prediction stated before evaluating six held-out datasets separated three MNIST-format datasets from three natural-photograph datasets by relative error reduction under both random-feature heads. A second test separated content from format: the gain on digits survives a change to color photographic images, but converting CIFAR-10 to grayscale also produces a gain, so image format contributes as well.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.