acceptodds
Under review as a conference paper at ICLR 2027

Not Order but Support: Curriculum Effects on Grokking Are a Transient-Support Phenomenon

Abstract

Grokking is when a network generalizes abruptly, long after it has already fit (memorized) its training set. It is sensitive to how the training data is presented—but “presentation” quietly bundles two different things, because a curriculum that shows hard examples first also shows them more often. We separate the two with a pair of controlled experiments. In the first, we fix the set of examples so that every example is seen equally often over the whole run, and vary only the order. Sorting by difficulty, in either direction, then prevents grokking outright on four of six small algorithmic tasks (validation accuracy stays at chance for the whole run even though the training set is fully memorized) and delays it 6–54× on the other two. So how often each example is seen is not what matters. In the second experiment, on two of the tasks, we go further and make every epoch (one full pass through the data) contain the whole difficulty range, so exposure is matched not just overall but at every epoch boundary, and again vary only the within-epoch order. The orders-of-magnitude delay then disappears: sorted and shuffled schedules reach the threshold within one evaluation interval. Thus, fine-grained order is not the primary lever either. The leading explanation is how long the optimizer is kept on a narrow, unrepresentative slice of the difficulty range. Curriculum effects on grokking are therefore a transient-support phenomenon; calling them ordering effects is imprecise. Two smaller results accompany this finding. First, a two-feature linear predictor built from early weight-matrix singular-value statistics forecasts the transition time with fully nested cross-validated between 0.66 and 0.81, requires no validation signal from the target run, and generalizes to held-out curricula. Second, at transformer scale, a sharp grokking-like transition appears on 2WikiMultiHopQA and reproduces across an 8× range of model sizes, but is absent or gradual on three other datasets. Because those experiments use an uncontrolled growing-pool curriculum, we present them as a boundary map rather than causal evidence about ordering. We release all 505 training runs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.