acceptodds
Under review as a conference paper at ICLR 2027

Why and When Deep is Better than Shallow: Implementation-Agnostic State-Transition Model of Deep Learning

Abstract

We ask when adding hidden layers improves generalization, in a model that keeps the layers fixed and varies only their number. A hidden layer is a self-map of a state space, a depth- network composes at most hidden layers with an output layer, and depth is compared within the nested family built from one class of hidden layers; this compares a deep network with shallower networks built from the same layers, not with wider ones. Our message is that the statistical cost of depth is the metric entropy of the set of compositions. The estimation error is bounded by an entropy integral over with constants that do not depend on the depth, and the bound is matched from below when the output layer can see the hidden states. For Lipschitz layers on a bounded state space this entropy grows at most polynomially in , and it stays bounded, or grows only like , under contraction, equicontinuity, or nilpotent structure. Balanced against the approximation error, this gives depths that grow with the sample size, and it separates the models whose estimation error is independent of the depth from those whose estimation error grows with it. Deep ReLU networks, unrolled solvers, and chain-of-thought computation are worked out; for the last two the number of steps is derived rather than assumed. All statements are machine-checked in Lean 4.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.