On the Exact Minimum Width for Universal Approximation with Deep ReLU Networks
Abstract
How narrow can a deep network be while still approximating every continuous map from to ? For ReLU networks, the exact minimum width has been known only in a few cases. We prove that it is exactly whenever , improving the previous upper bound by one, and show that the same holds for GELU, SiLU, Softplus, and other activations that approximate ReLU. Our proof builds on a recent diffeomorphism approach that separates the minimum width into a geometric quantity and the additional width needed to approximate diffeomorphisms. We show that one extra neuron suffices for every continuous nonaffine activation with a nonzero derivative at some point, giving the general upper bound . For ReLU, this extra neuron is necessary when approximating diffeomorphisms on full-dimensional regions, but universal approximation only requires approximation on the lower-dimensional image of . We exploit this distinction to remove the extra neuron when , where . For , we construct embeddings for which the geometric argument fails even after small perturbations. More generally, for all , the minimum width is either or , and we conjecture that it is always .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.