ResNets Are Deeper Than You Think
Abstract
The empirical success of residual connections is remarkable: nearly a decade after their introduction, they remain ubiquitous in modern neural network architectures. Their widespread adoption is usually attributed to improved trainability: residual networks train faster, more stably, and often achieve higher accuracy than their feedforward counterparts. While our experiments confirm the improved trainability of residual networks, we argue that this explanation is incomplete. We propose a complementary view: by combining shallow and deep paths in the computation graph of the network, residual connections can confer generalization benefits beyond pure optimization. Testing this hypothesis is surprisingly difficult: Simple side-by-side comparisons of residual and feedforward networks confound architectural prior with trainability, and prior attempts to close the gap do not isolate these effects conclusively. We therefore introduce a post-training partial-linearization setup that molds an already trained network into either a fixed-depth or a variable-depth architecture. In this setting, ordinary trainability differences are unlikely to explain the result. Across our experiments, variable-depth architectures consistently outperform fixed-depth counterparts at comparable nonlinear complexity, supporting the view that residual connections encode an inductive bias beyond numerical trainability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.