acceptodds
Under review as a conference paper at ICLR 2027

Data Schedules and Learning Rate Schedules Should Be Designed Together

Abstract

Mid-training (MT) or continual pretraining has become a common stage in modern large language model training pipelines. It follows the pretraining (PT) on web data, and includes smaller but higher-value data such as mathematics, code, reasoning traces, synthetic demonstrations, or long-context corpora. Although this practice is now widespread, it is still mostly treated as a practical technique rather than as a principal way of evolving training data distributions over time. In this work, we propose to view PT and MT as two phases of a unified data scheduling problem, in which the training distribution gradually shifts from broad web-scale coverage toward increasingly targeted objectives. This temporal choice is inherently coupled with optimization. The data schedule determines which distribution supplies the gradient at each point in training, while the learning rate (LR) schedule determines how strongly that gradient changes the model. Under this view, we reveal a critical but underexplored factor of data scheduling: the data schedule should be coordinated with the learning rate (LR) schedule. We find that while partially replaying PT data in MT substantially outperforms directly shifting data distribution from PT to MT data in the commonly used warmup-stable-decay LR schedule, its advantage diminishes when using a constant LR. Motivated by this contrast, we align the onset of LR decay with the increase in the MT data ratio. This matched decay range produces a better tradeoff between PT and MT loss than the other LR schedules we compare. Finally, we develop a tractable theoretical model that explains why the optimal allocation of a fixed MT data budget changes across analytically solvable LR schedules. Our results show that data and LR schedules should be designed jointly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.