LILA: Data-Free Structured Pruning of LLMs via NMF Spectral Geometry
Abstract
Structured pruning of large language models (LLMs) has universally required calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (Latent-Informed Layer Analysis) demonstrates that this dependency is avoidable: each FFN weight matrix is factorized with Non-negative Matrix Factorization (NMF) to obtain a parts-based spectral representation, and neuron importance is scored via the Kolmogorov Smirnov (KS) distance between singular value distributions of the full and neuron-ablated NMF matrix, a closed-form criterion requiring no calibration corpus, forward pass, or gradient computation. Two data-free matrix variants (Abs, Split) generate pruning masks from weight spectral geometry alone; a calibration-guided variant (Act) is provided as an internal ablation. Evaluation spans 7 model families (up to 70B). Without fine-tuning, data-free LILA surpasses PruneNet (45M-parameter RL policy) by up to 2.12 pp in zero-shot accuracy on LLaMA-2-7B and outperforms WikiText-2-calibrated SliceGPT by up to 6.1 pp across all sparsity levels, confirming data-free performance is not an artifact of omitting activation statistics. After one LoRA RFT epoch (matched WikiText-2 protocol), LILA Act surpasses SliceGPT+RFT by 6.39 pp on average using 98.4% fewer calibration samples (16 vs 1,024) across LLaMA-2-7B and Phi-2. NTK analysis confirms a reduction in functional distortion versus random pruning, providing theoretical grounding for the spectral criterion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.