OSD: Orthogonal Spectral Descent For Continual Learning In Large Language Models
Abstract
Continual fine-tuning enables large language models (LLMs) to adapt to evolving tasks and domains, but sequential adaptation often leads to catastrophic forgetting of previously acquired capabilities. Projection-based continual learning methods mitigate this interference by constraining updates to avoid directions associated with past tasks. However, these methods are typically built around Euclidean steepest-descent geometry, where update magnitude is controlled by the Frobenius norm. In parallel, recent matrix-aware optimizers such as Muon motivate a different view of neural-network optimization: for matrix-valued parameters, descent directions can instead be selected under a spectral-norm constraint. We introduce Orthogonal Spectral Descent (OSD), a practical continual-learning instantiation of spectral steepest descent that combines protected task subspaces with efficient Muon-compatible constrained updates. For selected matrix-valued parameters, OSD formulates each update as a spectral-norm-bounded descent problem subject to linear non-interference constraints derived from previously learned tasks. The resulting constrained problem admits a low-dimensional dual formulation that is approximated using a small number of warm-started dual updates together with Newton–Schulz matrix-sign iterations, avoiding an exact SVD at every optimizer step. We evaluate OSD across standard continual-learning benchmarks, TRACE, and domain-specific sequential fine-tuning settings using both encoder–decoder and decoder-only language models. Across the evaluated settings, OSD improves retention and overall continual-learning performance relative to sequential fine-tuning and competitive orthogonal-gradient baselines while maintaining strong adaptation to new tasks. These results suggest that combining protected task subspaces with spectral update geometry provides a practical way to improve the stability–plasticity tradeoff in continual LLM fine-tuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.