acceptodds
Under review as a conference paper at ICLR 2027

Curriculum Order Controls Grokking Through Its Hand-off Weights

Abstract

A neural network can generalize long after memorizing its training data; this late generalization is called grokking. A curriculum often helps by training the model on a chain of simpler tasks first, each starting from the weights the last one left. It is not known whether this chain passes on skills, or only a better initialization that another model could simply copy. We study a small transformer that solves a linear differential equation symbolically after a chain of negation, addition and reading off the solution's two coefficients. On the target alone, runs differing only in their random starting weights sometimes generalize and sometimes fail. Ordered easiest to hardest, the chain lifts average test accuracy from 0.48 to 0.88. And nearly every run now generalizes. But run in reverse order, ending on the easiest, the same tasks do no better than the target alone. To see which it is, we copy part of the chain's final weights into a fresh model trained only on the target. We split those weights in two: the token identity, which maps tokens in and out, and everything else. On its own, the token identity recovers 97% of the gain, although the fresh model has never seen the simpler tasks. And copying only the rest does worse than no curriculum at all. We repeat this on a second problem, symbolic differentiation, with its own chain, model and vocabulary. There the rest of the network carries the gain instead. So we find that a curriculum passes on an initialization, set by the last task before the target. The benefit of a curriculum is thus something one model can hand to another rather than earn by training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.