"Deep lifts": Neural networks can generalize even with the dumbest optimizer
Abstract
Given a labeled dataset , consider drawing a parameter vector uniformly at random from those that interpolate the data within a width- depth- ReLU architecture with continuous weights bounded within . We study the generalization properties of in the case where the dataset can, in fact, be explained by a smaller -by- ReLU network. Despite the obliviousness of the selection rule for , it adaptively generalizes at the rate , rather than being governed by the VC dimension of the full -by- hypothesis class. The reason is due to a particular inductive bias in ReLU networks: the -by- parameter witnesses an exponential volume of -by- interpolating parameters, which dwarfs the volume of poorly generalizing parameters (conditioned on interpolation). We explicitly construct this set of near-equivalent -by- parameters by a proof technique “deep lifting,” and then tie the achievable volume to stability properties of . Examples show that these techniques allow us to convert approximation-theoretic results into generalization guarantees for nonparametric classes. We further show that generalization of uniform-random interpolators can survive label corruptions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.