An Error Decomposition for Merging and Pruning Neural Networks
Abstract
While pruning removes units from a trained network, and merging combines several trained networks into one, many methods from both fields can be expressed as linear layer-wise compression methods. We introduce an error decomposition for such methods, which splits the layer-wise compression error exactly into a reconstruction error, which arises from feature information lost through compression, and a commutation error, which arises from the method's failure to commute with the network nonlinearities. Propagating these two local errors through the neural network gives a calibration-data based estimate of the output deviation between the compressed and initial network(s), which can be used to rank methods. Further the analysis reveals that the decoder—the linear operation that reconstructs the original features from the compressed representation—can be fitted independently in closed form once the encoder is fixed, because only the encoder is constrained to commute with the nonlinearity. Paired with the fitted decoder, a simple neuron-importance selection encoder, which has no commutation error, outperforms the considered merging baselines, including ZipIt and PLeaS, on same-task benchmarks across MLP, convolutional, and residual architectures. It performs on par with selecting the encoder using the predicted output deviation, showing that the decoder accounts for most of the downstream improvement. For Llama models, our experiments show that the derived reconstruction and calibration-selected axis allocation changes improve over FLAT-LLM constructions at matched parameter counts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.