acceptodds
Under review as a conference paper at ICLR 2027

Small LLMs: Pruning vs. Training from Scratch

Abstract

Pruning promises a shortcut to strong small models. However, Liu et al. (2019) found that in CNNs, this shortcut often loses to training from scratch. LLMs are substantially different, with larger model and data scales and different architectures, so the same conclusions may not hold. We examine this by pruning Llama-3.1-8B at pruning ratios of 0.5–0.8 with six methods spanning depth, width, and sparse granularities, under two token-matched settings. (1) With the same training token budget, pruned initialization consistently outperforms random initialization. This shows that the parent model provides a strong starting point, although the advantage narrows as the training token budget grows and, for depth pruning, as the pruning ratio rises, nearly vanishing at the highest ratio we study. (2) When training from scratch is instead given the full token budget consumed by the whole pipeline, pruning at finer granularities still retains a small advantage, while coarser structured pruning can be matched or surpassed. This suggests that the parent model may transfer knowledge that additional tokens alone cannot fully recover, but only at fine granularity. This forms a clear recommendation: with a strong parent and a limited token budget, pruning is more effective; when tokens are plentiful, training from scratch is competitive at coarse granularity, so a parent is not always necessary.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.