acceptodds
Under review as a conference paper at ICLR 2027

Theory of Feature Formation, Expression and Reorganization in Quadratic Networks

Abstract

A neural network learns a task by forming useful features and learning how strongly to use them in its predictions. Classical exact mode solutions for deep linear networks describe how expression grows from weights already aligned with the task, leaving the formation of those features unaddressed. We derive how feature directions form from unaligned weights in a quadratic network and how growing strengths affect further direction learning. Under population gradient flow, we solve the early direction dynamics in the limit of small initialization across a broad class of input distributions. The network acquires a hierarchy of directions set by the task and data while its output remains small, with alignment timescales governed by differences in growth rates. For isotropic Gaussian inputs, the direction law is exact throughout training and does not require small initialization. We quantify silent alignment feature by feature, showing that smaller initialization delays substantial use of each feature without slowing its formation. For more general inputs, growing feature strengths can feed back on direction learning through the changing prediction error and reorganize the features formed early in training. Exact solutions for structured input distributions show how anisotropy and fourth moments govern this feedback-induced reorganization: the network can replace a feature acquired early or mix it with another direction, or features can promote or suppress one another's growth. The resulting theory connects the formation of a representation to its subsequent development as the network learns to use it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.