Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules
Abstract
LLMs encode both general capabilities and domain-specific knowledge in a single set of parameters. We ask whether this capacity can be reorganized: keeping broadly useful computation in a shared backbone, while moving specialized knowledge into external memory modules. We propose knowledge offloading (KOFF), a framework for decomposing a pretrained LLM into a sparse shared backbone and domain-specific memories. Starting from a frozen base model, we either jointly or sequentially learn a structured pruning mask and lightweight recovery modules, implemented as LoRA adapters and learned key-value caches. Across Llama and Qwen models from 3B to 14B, we find that 12% to 20% of capacity can be moved out of the shared backbone while largely preserving perplexity and MMLU performance (up to 90% of the full model). We show that LoRA and learned KV memories are complementary, and analyses suggest that the offloaded knowledge is specialized: mismatched memories do not yield the same gains, and language-specific neurons are preferentially removed while language-neutral neurons largely remain in the backbone. These results suggest that the knowledge of a pretrained LLM can be reallocated between a shared core and swappable external memories.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.