acceptodds
Under review as a conference paper at ICLR 2027

Data Geometry and the Metric Entropy of Infinite-Width ReLU Networks

Abstract

The complexity of an infinite-width network is measured in , so it is a property of the data distribution and not of the architecture alone. For ReLU ridge features that dependence runs through a single quantity: the smoothness of the map from parameters to features, which is fixed by how much probability mass sits near moving activation boundaries. An anti-concentration, or margin, exponent for those boundaries therefore determines the metric entropy of the model class. We first show that this exponent can never exceed one, whatever the data, so a margin law imposed uniformly over the parameter space can only degrade the classical smoothness of ReLU features and never improve it. We then show that the degradation it predicts can be an artifact of that uniformity, and is so in the cases that motivate the question. Our main result is a stratified entropy law, in which the parameter smoothness is allowed to fall off on a set of positive codimension at a controlled rate; the resulting entropy exponent is governed by the better of two competing terms, and reduces to the smooth-dictionary bound of Siegel and Xu (2024) w when the bad set is everything. The motivating case is data concentrated on a curved low-dimensional manifold, where the margin exponent does drop, but only at parameters whose activation boundary is tangent to the data — a set of codimension one. There the stratified law returns the undegraded classical exponent, reconciling the margin mechanism with intrinsic-dimension theory, and we delimit when a margin exponent below one arises at all. The predicted local exponents, and the rate at which the smoothness degrades near the bad set, are confirmed numerically to four digits.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.