TraceNAS: Neural Architecture Search for LLM Pruning via Gradient Trace Correlation
Abstract
Structured pruning is essential for efficient deployment of Large Language Models (LLMs). The varying sensitivity of LLM sub-blocks to pruning necessitates the identification of optimal non-uniformly pruned models. Existing methods evaluate the importance of layers, attention heads, or weight channels in isolation, ignoring the complex global structural dependencies that exist across the model. Training aware search for structured pruning addresses these dependencies by fine-tuning candidates during the search, making the search nearly as expensive as the recovery training that follows. We propose TraceNAS, a Neural Architecture Search (NAS) framework for joint depth and width structured pruning that replaces this with a training free search: each candidate is scored without any weight updates, using a scale-invariant zero-shot proxy whose alignment with the pretrained model correlates with its potential to recover accuracy during post-pruning training. By removing weight updates from the search loop, TraceNAS discovers high-potential pruned architectures on a single A100 GPU in 8.5 hours, yielding up to 7.5 reduction in search wall clock, 1.35 reduction in search peak memory, and 10 fewer tokens processed during search compared to training-based search baselines. Evaluations on the Llama and Qwen families show TraceNAS matches training aware search performance at a fraction of the search cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.