The Model Shape Behind Scaling Laws: A Gradient Superposition Perspective
Abstract
Neural scaling laws guide resource allocation, but how should a fixed parameter budget be divided between depth and width? Motivated by inverse-width representation-error scaling under strong superposition, we study this allocation through the network's end-to-end response. The Neural Feature Ansatz links weight and gradient geometry, motivating the Frobenius norm of the input average gradient outer product (AGOP) as an observable of response strength and overlap. By reproducing Anthropic's toy-model double-descent experiment, we find that and test reconstruction loss follow an approximately affine relationship across the full sample-size sweep and share an interpolation peak. Through Gaussian minimum-norm regression experiments, we further show that this correlation arises from a fitted-noise amplification term shared by prediction error and . By pretraining language models spanning different shapes and parameter budgets, we find a strong positive correlation between loss and among the best shapes selected by validation loss. This supports interpreting effective parameter allocation as preserving task-relevant computation while reducing the amplification of error responses. Our results identify error-response amplification as a mechanism linking gradient superposition to generalization, providing a functional account of depth-width allocation and efficient parameter utilization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.