acceptodds
Under review as a conference paper at ICLR 2027

Structural re-parameterization is not magic

Abstract

Some vision models are trained with several parallel kernels per layer, each with its own BatchNorm, summed during training and folded into a single kernel at inference, a technique known as structural re-parameterization. It has been reported to give a modest accuracy benefit, explained so far by implicit ensembles, diversity and "training-time overparameterization". We study a simple instance, two identical branches, and take it apart. We attribute the accuracy benefit to two simple mechanisms, previously unidentified. The mechanisms are optimization effects: a coordinate specific learning rate boost at the last layer and a directional preconditioning effect at the first layer. The boost acts on the scale coordinate and helps to grow the logit scale at the end of training. The preconditioning cools the learning rate of the highest variance mode of the input, making the system better conditioned. We show this by dissecting the model one layer at a time, deriving the modified step in closed form, and seeing that applying the corresponding optimization modifications to a single branch reproduces the observed benefit. We conduct our study in a small VGG like network on CIFAR and then show that the two mechanisms carry over to production architectures on ImageNet.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.