Permutation-Packable N:M Sparsity for Sparse-to-sparse Muon
Abstract
Applying Muon to sparse vision Transformers creates a mismatch: its Newton–Schulz (NS) iterations can generate updates at pruned weights, which projection then discards, while dense training tensors can limit the memory benefits of sparsity. We introduce Permutation-Packable N:M (PP-N:M), a locally balanced support that is N:M in both directions and packs exactly into dense active matrices. Its closure under NS cubic products enables matrix updates entirely on active coordinates. On this support, S2S Muon constructs update directions through pre-NS row–column equilibration and component-wise NS, with a second-moment-based Adam reference controlling their magnitude. To retain compact storage beyond the optimizer, we develop PP-linear kernels that fuse channel indexing with matrix products and compute packed weight gradients directly. In a seven-method ImageNet-1K comparison at 300 epochs over two seeds, ViT with S2S Muon reaches 78.21% mean top-1 accuracy versus 77.31% for projected Muon, a gain of 0.90 percentage points, while retaining half of the eligible weights. Complementary CIFAR-100 experiments reach % (mean standard error) across three seeds, 3.76 percentage points above projected Muon. As a cross-domain test, 350M and 1B language models improve perplexity over projected Muon by 2.59% and 2.31%, respectively, at a fixed 1.008B-token budget. On RTX 5090, PP-linear reduces peak allocated training memory for single-momentum Muon+PP-N:M by 18.29% on ViT and 40.22% on a 1B language model relative to projected Muon. With the complete S2S Muon update rule held fixed, it saves 15.17% and 29.09%, respectively, relative to a dense-masked implementation. These memory benefits come with workload- and implementation-dependent throughput trade-offs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.