MemSFT: Mitigating Alignment Tax with an External Parametric Memory
Abstract
Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastrophic forgetting on general tasks. We propose MemSFT, which mitigates the alignment tax by decoupling domain specialization from backbone LLM parameter updates through a plug-and-play parametric memory. Building on Memory Decoder (Cao et al., 2026a), which primarily studies external memory for pretrained LLMs, we develop supervised fine-tuning on external memory to support domain specialization of post-trained LLMs. To selectively invoke domain expertise, we introduce a dynamic memory injector that fuses memory and backbone output distributions at each decoding step. We evaluate MemSFT across three domains, namely biology, geoscience, and law, covering small dense and large mixture-of-experts (MoE) LLMs ranging from 8B to 397B parameters. MemSFT consistently improves domain performance without catastrophic forgetting of general capabilities, whereas full SFT suffers severe forgetting on general tasks. Our results demonstrate a practical path to domain specialization of modern LLMs, enabling memory reuse across multiple backbones with only 0.27× and 0.41× the adaptation FLOPs of full SFT and LoRA for a 397B model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.