Emergent, Coordinate-Specific Structure in Dense Parameter Updates for Continual Learning
Abstract
Continual learning (CL) requires models to learn new tasks while retaining previous knowledge, yet sequential parameter updates can cause catastrophic forgetting. Parameter-efficient adaptation methods such as LoRA address this challenge by restricting updates to a low-rank subspace, limiting the degrees of freedom of each task update. We ask the converse: what structure emerges when the full parameter space remains unconstrained and is updated continuously across tasks? We freeze a pretrained model and train a same-sized parameter state from zero with standard AdamW. We find that continuous dense updating develops cross-task co-directionality and a contracting effective rank, in contrast to the rank expansion observed in LoRA, and that this structure favours knowledge retention across tasks. Coordinate and trajectory interventions show that retention depends on both where updates are written and whether their accumulation remains continuous: preserving update values while permuting their coordinates sharply reduces retention, while interrupting the trajectory at task boundaries also degrades it. At matched learning amounts, LoRA incurs greater per-task parameter displacement, accounting for much of its additional forgetting, with residual dependence on rank. Across the swept range, the dense adapter achieves better retention than LoRA, while its frontier intersects that of EWC. These results reveal that continuous unconstrained updating itself forms structured parameter dynamics with measurable consequences for continual learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.