acceptodds
Under review as a conference paper at ICLR 2027

Weight-Scale Growth Tracks Training Effort, Not Learning Quality

Abstract

Does weight growth mean learning? Previous work showed that data predictability forecasts growth in the Weibull weight-scale parameter λ during transformer training. We test what this growth records, why it can diverge from generalization, and how it accumulates. In a controlled factorial grid crossing token shuffling with fixed-budget repetition, fully shuffled, highly repeated data produces the largest Δλ² but the worst held-out loss. Under fixed AdamW, Δλ² closely tracks accumulated radial-alignment work. A component- and time-resolved follow-up further shows that data conditions change both when this effort accumulates and whether it is routed toward Transmission or query/key Selection weights. Within the main grid, pairing Δλ² with the train–held generalization gap separates learning, memorization, and nothing-to-learn without seed-level label flips. The same frozen thresholds preserve this separation under matched corner interventions on C4 and Python code. These experiments trace weight-scale growth from controlled data construction through optimizer work and component-resolved weight evolution to behavior. Weight-scale growth therefore describes training activity, but alone does not determine whether the fitted structure generalizes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.