acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Catastrophic Forgetting in LLM Continual Learning through Attractor-Informed Latent Transition Alignment

Abstract

Continual learning enables large language models to acquire new knowledge while preserving the computational structures that support their existing capabilities. In replay-free continual learning, however, original pre-training data are unavailable, and model updates may disrupt these structures and cause catastrophic forgetting. Existing methods mainly preserve model parameters, output distributions, or isolated hidden states, while paying less attention to cross-layer transformations. We propose Latent Transition Alignment (LTA), which uses cross-layer latent transitions as a preservation target. In selected intermediate Transformer windows, our latent-dynamics analysis identifies coherent transition fields and locally attracting fixed points. We then introduce an offline dual-anchor surrogate that connects observed transitions to the residual field of a latent self-map. Under exact reconstruction, the observed latent transition equals the residual-field vector at its start state. Under approximate reconstruction, their discrepancy is bounded. Online, LTA aligns transition directions and applies one-sided magnitude control without online field fitting or fixed-point iteration. In single-domain learning, LTA outperforms all evaluated baselines in retention while matching SFT in new-domain learning; relative to Output KL, it reduces accumulated first-domain forgetting by 8.20 percentage points and improves final QASC accuracy by 7.30 percentage points in multi-domain continual learning; under fixed-protocol transfer to Qwen2.5-3B, it requires no scale-specific retuning and reduces forgetting by 63.1% while improving QASC accuracy by 3.33 percentage points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.