acceptodds
Under review as a conference paper at ICLR 2027

Curriculum Learning: Does Data Order Matter?

Abstract

Curriculum learning, training on data in a defined order rather than at random, is one of the oldest ideas in machine learning, yet its gains in language-model pretraining have been unreliable. In this work, we ask whether data order matters at all. Holding the data mix fixed, we find that the same curriculum can help or hurt depending on the corpus. Compression-ratio interleaving improves GSM8K and MATH-500 by 5–7 percentage points (pp) on mathematical web text, but on general web text reduces ARC-Easy by 3–7 pp and GSM8K by 6–8 pp. These results suggest that curriculum effectiveness depends on whether an ordering aligns with useful structure in the corpus. We therefore introduce such structure during data generation, rewriting mathematical problems under grade-specific pedagogical rubrics. When a 1.7B model sees these 10B tokens of synthetic math, presenting grades 1 through 7 in pedagogical order raises GSM8K and MATH-500 by 4 pp over randomly sampling that data. Our results suggest that when natural data lacks a useful order, synthetic rewriting offers a general way to create one.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.