From Reasoning to Agents: Building the Next Generation of On-Device Language Models
Abstract
Recent advances in agentic models have enabled a growing range of applications, driving increasing demand for training models specialized for downstream agentic tasks. Existing work, however, either develops proprietary construction recipes for large-scale models or optimizes post-training in isolation, leaving the end-to-end training of small agentic models poorly understood. We show that downstream agentic performance is largely determined by the capabilities and behavioral support established during pre-training; post-training sharpens what the base model can already express, but does not reliably recover what it cannot. We therefore ask which pre-training signals predict post-training gains. We find that sampling lift, the improvement obtainable by sampling multiple rollouts over greedy decoding, serves as a good indicator, suggesting that pre-training should preserve diverse behavioral support rather than prematurely concentrating on narrow solution modes. Since pre-training metrics differ substantially in predictive power, we weight them by their measured correlation with post-training outcomes, providing a principled basis for checkpoint selection and data-mixture allocation, and allowing pre-training to be concentrated on the capabilities most predictive of downstream performance. Building on these principles, we study how capabilities equipped in general pre-training are converted into reliable agentic behavior. We find that long-horizon reasoning does not transfer directly to multi-turn tool use, while separating the decoding pattern and injecting trajectory-level supervision enables more effective learning from the student's evolving state. Together, these results form an end-to-end recipe for small agentic models. Trained from scratch with only 1.2T tokens, our 2.5B model matches or surpasses the 7.3B-parameter OLMo-3 model trained with 5x more tokens, achieving the best overall results on AIME24, GSM8K, HumanEval, and SWE-Bench.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.