acceptodds
Under review as a conference paper at ICLR 2027

From Class Means to Margins: Predicting Hidden-Unit Placement in Width-One Networks

Abstract

At small output scale, logistic loss favors separation of class means; at large scale, it favors the worst-case margin. We study a width-one network with activation on symmetric inner-versus-outer intervals, the smallest setting in which these preferences require different hidden features. An independently re-checked certificate locates the placement switch, and a scaling reduction predicts its activation dependence with a first-order coefficient confirmed at finite parameter values. Registered prospective tests predict training crossings from conditional thresholds fixed before training. A first-order tracking law, derived post hoc, predicts each run's lag from the branch geometry and its measured ratio of output-scale growth to hidden relaxation. At an activation value with no prior crossing data, its registered test gives median per-run observed/predicted ratios of 1.07 under Adam and 1.06 under SGD; individual predictions are accurate under SGD, where every run is within 10%, but not under Adam, where 22% are. A registered learning-rate test finds that the lag does not vanish as the learning rate decreases. This quantitative predictive link is established for width-one sine networks that reach a branch and grow slowly relative to relaxation; outside this setting, registered predictions fail or remain unresolved, and agreement conditioned on the eventual branch is explanatory.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.