RASA: Retention-Aware Stream Adaptation for Online Continual Learning in Large Language Models
Abstract
Continual learning (CL) enables large language models (LLMs) to learn new tasks while preserving previously acquired knowledge. Expansion-based methods extend the model through task-specific modules, with low-rank adaptation (LoRA) enabling parameter-efficient expansion. However, in online continual learning (OCL) settings where data arrive in a stream and cannot be revisited, current-task information alone is insufficient to evaluate how adaptation affects performance on previous tasks. In this work, we propose Retention-Aware Stream Adaptation (RASA), a probabilistic framework that combines online statistics of parameter changes with historical feature information. RASA models parameter changes induced by the optimization process as a local Gaussian transition distribution, whose covariance is matched to a reference covariance estimated from the stream to determine adaptive time steps. To incorporate retention, it uses aggregated feature second moments from previous tasks to constrain cumulative changes in LoRA linear outputs on historical input features. A local generalized posterior then combines this constraint with gradient information from the current batch for parameter estimation. The resulting procedure can be incorporated into existing training procedures, reusing available gradients and stored feature statistics without replay or additional forward or backward passes over the data. Experimental results demonstrate that RASA consistently improves the average performance of existing expansion-based methods in OCL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.