acceptodds
Under review as a conference paper at ICLR 2027

Strong Backbone, Simple Heads: The Case for Gradient-Free Class-Incremental Learning

Abstract

Recent class-incremental learning methods increasingly build on pre-trained vi- sion transformers, adapting only a small set of parameters such as lightweight adapters. However, these methods are typically developed and benchmarked on older foundation models. We ask how they transfer to modern backbones, and whether parameter adaptation is needed at all. Comparing adapter-based meth- ods with gradient-free, closed-form classifiers that keep only per-class statistics, under matched backbones, splits and seeds, we find that as backbones improve the gradient-free methods match or outperform gradient-based ones, even after extensive retuning of the latter. The gap is driven by forgetting rather than learning: the adapter method learns each task about as well as the closed-form heads, but forgets roughly three times as much, and its forgetting is higher on the stronger backbones while theirs falls. We also find that the composable, sufficient-statistic state of these heads enables two practical properties: federated merging across distributed instances and class-level machine unlearning. These results suggest that with strong foundation models, simple statistical decision rules can outperform increasingly elaborate adaptation mechanisms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.