acceptodds
Under review as a conference paper at ICLR 2027

Which neural network layers have stable representations?

Abstract

Many applications that rely on deep neural network (DNN) representations, such as encoding models or transfer learning, require choosing which intermediate layers to use. For such applications, we would like to get similar results for different instances of a model, i.e. stable results. Here, we aim to find out which DNN layers fulfill this criterion. We analyze five instances each of six ImageNet-trained architectures, comparing representations from each layer with their counterparts in other instances. Later network stages tend to be less self-similar, while certain layer types maintain higher stability throughout the network. We identify three criteria that characterize these most stable layers: (i) they are not bypassed by a skip connection, (ii) they are normalized, and (iii) they are placed after the non-linear activation function. Based on these criteria we modify ConvNext to yield an additional stable layer type confirming the validity of our criteria. Furthermore, we predict a large-scale fMRI dataset from each layer's representation, finding that stable layers yield more consistent and improved results for these encoding models. Finally, we tested transfer learning for a fine grained classification task, again observing more stable and better performance for stable layers. To confirm that our observations are not caused by the specific comparison methods, we perform each analysis with two different sets of methods. Based on our results, it seems advisable to focus on stable layers for all analyses that are based on internal representations of DNNs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.