acceptodds
Under review as a conference paper at ICLR 2027

Can Old Dogs Learn New Tricks? Capability Dynamics Along Pretraining

Abstract

When teams continue training a language model, they have to choose which data to train on, which checkpoint to start from, and how much old data to mix back in. These choices are usually guided by the training loss and by benchmark scores based on answer likelihood. We test whether these signals lead to the right choices. We take public OLMo-2 checkpoints and continue each one several times. Each run changes a single setting and keeps all training batches identical. We judge the model by the answers it writes. Likelihood scores can pick the wrong data. Training a 1B model on math data raises its accuracy on word problems by 47 percentage points while code data does not, yet a likelihood score computed on the final answer alone, with no written reasoning, favors code. The starting checkpoint also matters. At 7B, with our fine-tuning recipe, checkpoints from before mid- training learn a new task far more slowly than later ones. Finally, runs that look the same on loss and benchmark averages differ in which questions they get right. On GSM8K the difference is mostly a missing stop token.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.