acceptodds
Under review as a conference paper at ICLR 2027

Understanding the Role of Reasoning Data in Pretraining

Abstract

Reasoning data expose models to multi-step reasoning, but a reasoning trace may reveal only a subset of the underlying facts it composes. How should pretraining balance atomic facts with reasoning traces? We study this question by constructing a setting in which every fact provides an atomic training signal, while reasoning traces contain only a subset of the facts they compose, enabling us to measure generalization to unseen transitions. For a linear transformer, we derive the weight assigned to atomic training signals that minimizes held-transition squared error and show that its leading-order optimum increases with reasoning-chain length and decreases with the fraction of facts held out from reasoning traces. These trends persist in a full transformer layer with an MLP, softmax attention, and cross-entropy loss, as well as in a realistic mathematical reasoning setting. Together, our results provide a principled view of how reasoning-chain structure should shape the allocation of training data during pretraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.