acceptodds
Under review as a conference paper at ICLR 2027

Layer-Wise Input Recoverability in Deep Neural Networks

Abstract

The inner workings of deep neural networks are still not well understood. The information bottleneck framework, which characterizes deep networks by the mu- tual information between their hidden representations and the network’s input, has been used to identify compression of input information with respect to an input image and output labels. We instead study how much of a layer’s own input is recoverable from the layer’s output in convolutional neural networks (CNNs). A convolutional layer applies its filters to one input patch at a time, so we eval- uate recoverability on sampled input patches with a reconstruction-based probe and not on the entire feature map. The probe’s size thus depends on channels and kernel size and not on feature map size, which lets our method scale to deep CNNs with wide feature maps. For each layer, we compare pretrained, random, and reconstruction-optimized encoders, using linear and linear-plus-residual-MLP decoders. We measure recoverability as R2 between input patches and their re- constructions on a held-out set. Across nine ImageNet CNNs, recoverability gen- erally decreases with depth but is more consistently predicted by the compression of the layer. Learning-attributed recoverability, the difference between pretrained and random encoders, is not uniform across architectures. We introduce single- property random controls, which transfer one pretrained property onto a random encoder, and find that each control changes recoverability with a consistent sign across models: copying pretrained singular vectors tends to raise recoverability, while copying the pretrained singular-value spectrum or activation sparsity lowers it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.