Normalized Factorized Weight Generation for Nested Elastic Convolutional Networks
Abstract
How should a weight generator be parameterised so that one stored representation supports nested neural networks without making shared weights depend on the selected width or depth? We propose Spectrum-Shared Generation (SSG), a normalized factorized parameterisation for elastic convolutional networks. Each residual kernel is a CP map with layer-specific directions and network-wide component magnitudes. Normalising complete factor columns before channel selection makes every narrower generated kernel an exact prefix sub-tensor of wider ones; fixed layer identities similarly preserve the generated tensor assigned to a retained block under depth selection. The parameterisation is invariant to radial factor rescaling and admits explicit local partial-Jacobian bounds, while configuration-specific normalisers handle activation statistics outside the generator. In controlled comparisons, SSG reduces persistent learnable storage by – across CIFAR-10, CIFAR-100, and Tiny-ImageNet, with mean accuracy differences of – points from dense elastic supernets. On CIFAR-100, one 2.084M model serves eight configurations spanning in compute and reaches 76.67% at full capacity, 1.14 points above an OFA baseline retrained under our protocol while using less persistent learnable storage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.