acceptodds
Under review as a conference paper at ICLR 2027

What Should Continual LoRA Remember? Historical Activation Geometry for Fixed-Capacity Adaptation

Abstract

Continual LoRA updates a shared low-rank adapter across tasks, but preserving previous behavior under fixed capacity depends on what information is retained from past data. We first show that adapter updates alone cannot identify their historical layer-response drift uniformly over past input distributions: the same update can induce zero or maximal drift on different histories. Guided by this result, we compare two uses of few-shot history, replaying labeled examples in the training loss and encoding prompt-only examples as activation geometry. We propose HARD AG-LORA, which stores a capped reservoir of historical layer activations, freezes the LoRA output factor, and projects subsequent input-factor displacements away from the leading historical activation subspace. This preserves direct adapter responses on the protected coordinates while maintaining a single fixed-rank adapter; a controlled variant further shows that storing activations when tasks are learned is more effective than recomputing them under the current model. Across three orders of four text-classification tasks on Qwen2.5-7B, HARD AG- LORA achieves 80.96% final average accuracy with 0.95 points of forgetting, compared with 77.82%/5.36 for Replay-64 and 80.72%/1.20 for growing-rank O-LoRA. On Meta-Llama-3.1-8B, it achieves 78.51%/4.58 versus 75.14%/8.20 for rank-32 O-LoRA. HARD AG-LORA keeps one rank-8 adapter and a task-count- independent activation sketch, providing an actionable fixed-capacity representation of historical information for continual adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.