Sparse Neural Thickets: Sampling the Small Weights Subspace
Abstract
The loss landscape around pretrained models contains dense neighborhoods of task-improving solutions, known as _Neural Thickets_. We show that this structure persists even when perturbations are restricted to low-dimensional parameter subspaces, revealing _Sparse Neural Thickets_. Strikingly, subspaces of low-magnitude parameters are consistently enriched with high-performing solutions compared to subspaces of large-magnitude parameters and dense subspaces. This suggests that favorable regions of the local loss landscape are concentrated in an easily identifiable subset of parameters. Leveraging this structure, we introduce **Su**bspace **P**arameter-**E**fficient **S**ampling (SuPES). The higher density of good solutions enables more efficient exploration: SuPES achieves higher accuracy at comparable sampling budgets or comparable accuracy with fewer samples, while perturbing substantially fewer parameters than full-space sampling. Experiments across language models, parameter scales, and datasets establish Sparse Neural Thickets and demonstrate their potential for parameter- and sample-efficient sampling around pretrained models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.