Downsampling Blocks and Residual Connections Propagate Memorization Entangled with Generalization in ResNets
Abstract
Residual networks (ResNets) are among the most widely used deep learning models due to their ability to train very deep neural networks through a stable gradient flow enabled by residual/skip connections. However, no prior work has studied how different model pathways propagate memorization and generalization within the ResNet architecture. Hence, in this work, we investigate which pathways in ResNets are responsible for propagating memorization and generalization. We find that downsampling blocks and residual connections propagate memorization along with generalization, supported by activation-level analysis. We conduct a block-wise Fisher information analysis to elucidate the dominant role played by downsampling blocks in propagating memorization. We also provide a weight-level analysis to show that weights contributing to memorization and generalization are entangled and primarily concentrated in downsampling blocks. We validate our findings across three interpretability strategies, six architectures, and six datasets. Together, our results provide a clear map of how memorization and generalization propagate in the ResNet architecture, which has never been articulated in the literature.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.