acceptodds
Under review as a conference paper at ICLR 2027

DIMENSIONALITY EXPANSION IMPROVES SPARSE FEATURE LEARNING IN NEURAL NETWORKS

Abstract

Sparse dictionary learning is the standard tool for decomposing neural network activations into interpretable features, and every deployed pipeline must choose an expansion factor E, the ratio of dictionary width to activation width. That choice is made by convention rather than by measurement, because on real activations the ground-truth features are unknown and recovery cannot be scored. We have built 1024 features in superposition in R64 under a heavy-tailed frequency law. Across 153 training runs, recovery rises monotonically with expansion, from mean max cosine similarity 0.374 at E = 1 to 0.703 at E = 32, a 9.7× increase in the fraction of features isolated at a cosine threshold of 0.9. Over the same range normalised MSE improves by less than half, so tuning width by reconstruction error substantially understates what expansion buys. The relationship is affine in log2 E (R2 = 0.981) and a random-allocation null model is rejected. The benefit is strongly regime-dependent: it is largest when few features are co-active (0.457 MMCS gained at k = 2) and nearly vanishes when the support is dense (0.092 at k = 32), so width cannot rescue a problem that has left the identifiable regime. Finally we introduce Progressive Dimensional Expansion (PDE), which grows into a target width while recycling duplicated and dead capacity. At matched compute PDE improves MMCS over fixed-width training by 0.0073 (paired t(9) = 4.98, p = 0.0008) and recovery rate by 15.0% relative, using 95% of the floating point operations. An ablation shows the gain comes from capacity recycling rather than from the growth schedule

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.