acceptodds
Under review as a conference paper at ICLR 2027

Base Models Can Reason by Taking a Cue from Their Training Data

Abstract

In this paper, we study how training data creates associations between the first tokens a base model outputs and its reasoning behaviors. First, we discover that fixing particular starting *token cues* makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, we demonstrate that RL itself learns to make these same starting cues more likely, while fixing these cues recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal interventions on the data to turn an arbitrary word, such as "Chicken", into an effective reasoning cue or remove an existing cue's effect. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues to a safety setting, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.