acceptodds
Under review as a conference paper at ICLR 2027

The Model Shape Behind Scaling Laws: A Gradient Superposition Perspective

Abstract

Neural scaling laws guide resource allocation, but how should a fixed parameter budget be divided between depth and width? Motivated by inverse-width representation-error scaling under strong superposition, we study this allocation through the network's end-to-end response. The Neural Feature Ansatz links weight and gradient geometry, motivating the Frobenius norm of the input average gradient outer product (AGOP) as an observable of response strength and overlap. By reproducing Anthropic's toy-model double-descent experiment, we find that and test reconstruction loss follow an approximately affine relationship across the full sample-size sweep and share an interpolation peak. Through Gaussian minimum-norm regression experiments, we further show that this correlation arises from a fitted-noise amplification term shared by prediction error and . By pretraining language models spanning different shapes and parameter budgets, we find a strong positive correlation between loss and among the best shapes selected by validation loss. This supports interpreting effective parameter allocation as preserving task-relevant computation while reducing the amplification of error responses. Our results identify error-response amplification as a mechanism linking gradient superposition to generalization, providing a functional account of depth-width allocation and efficient parameter utilization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.