A FEW GOOD WEIGHTS: SPARSE SUBNETWORKS IN PREFERENCE-OPTIMIZED LLMS
Abstract
Reinforcement learning fine-tuning of Large Language Models (LLMs) induces sparse, full-rank weight updates concentrated subnetworks containing fewer than of parameters. We show the structure of these subnetworks is determined primarily by the learning objective, rather than the training data. Subnetworks produced by DPO and GRPO differ substantially in parameter selection and internal representation, even on similar data, while subnetworks produced by the same objective on different datasets remain closely aligned. To establish this result rigorously, we develop a theoretical framework for identifying the sparse subnetwork by scoring each parameter's importance. We derive an optimal Oracle scoring function with guarantees on loss retention, prove the existence of an optimal scoring function at any sparsity level, establish a certifiability condition for evaluating practical scoring functions, and develop a scalable threshold estimation algorithm for constructing subnetworks in multi-billion-parameter LLMs. We propose warm-start magnitude scoring, which recovers dense DPO and GRPO training dynamics at or beyond sparsity across mathematics, coding, syntax, and semantics. We characterize objective-specific subnetwork divergence through Jaccard similarity, Centered Kernel Alignment, and linear probing. Finally, we introduce a sparsity-aware AdamW optimizer kernel implemented in Triton that reduces optimizer step complexity by the factor , achieving and speedups over PyTorch and HuggingFace 8-bit AdamW respectively at sparsity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.