Learning While Forgetting: Sparse Adapter Consolidation for Memory-Bounded Continual Learning
Abstract
Continual learning under a fixed retained-state budget eventually forces a choice about what to forget. Retaining one adaptation per task grows storage with the stream, while consolidating adaptations into shared memory makes older information progressively harder to recover. We introduce Palimpsest, a write–recover–gate learner that treats consolidation and forgetting as coupled problems. Each task's low-rank adapter is sparsely represented and superposed, with decay, into a shared fixed-size history store. Joint sparse recovery reconstructs candidate historical adapters. A support-recovery analysis motivates an age-dependent reliability horizon, while a calibrated load correction turns this into a deployment forecast that determines when to serve a recovered adapter and when to fall back to the frozen backbone. Reusing the ImageNet-R calibration without refitting predicts the recovery horizon of a running class-incremental learner with . Gate interventions show that fallback recovers 3.1–14.1 percentage points of mean old-task accuracy after the predicted horizon under forced forgetting, while also exposing conservative cases in which useful adaptation is discarded. Under retained-state caps of 5M parameters on ImageNet-R and 15M on OmniBench-1K, Palimpsest leads both reported accuracy metrics at the longest evaluated horizons; at 200 tasks on OmniBench-1K, it exceeds the strongest capped baseline by 7.18 points in average accuracy and 6.90 points in final accuracy. The advantage emerges as memory becomes limiting, supporting sparse consolidation with explicit, forecast-guided forgetting as a principled approach to memory-bounded continual learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.