acceptodds
Under review as a conference paper at ICLR 2027

Towards 1:4 semi-structured 75% sparsity via Cannistraci-Hebb N:M dynamic sparse-to-sparse training

Abstract

Improving the computational and energy efficiency of large language models motivates semi-structured N:M sparsity, but preserving model quality becomes difficult when moving from 2:4 (50%) to 1:4 (75%) sparsity. We propose CHTsNM, a sparse-to-sparse pretraining framework that combines Cannistraci–Hebb topology evolution, topology-aware optimization, and contextually modulated low-rank compensation. CHTsNM directly trains the executed sparse backbone while evolving legal row-wise or transposable block-wise supports. Its Topology-Aware Nesterov Optimizer (TANO) stores moments on active coordinates, retains the history of surviving connections, and resets newly activated states. Contextually Modulated LoRA (CoMoLoRA) supplies input-adaptive residual directions through rank-space modulation. These components address connectivity, optimization, and the capacity lost under a sparse constraint. In the completed two-seed LLaMA main study, CoMoLoRA reduces mean perplexity in all twelve auxiliary/no-auxiliary configuration pairs, with larger gains at 1:4 than at 2:4. At 350M and 1:4, the gains are 2.124 and 2.423 perplexity points for row-wise and block-wise supports. Matched 130M controls show complementary benefits from TANO and residual compensation, while packed FP32 moments reduce eligible-matrix state storage by 75%. These results support combining effective optimization, compact optimizer state, and residual compensation as a practical route toward high-quality 1:4 sparse pretraining.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.