Davinci-Math: A Routing-Aware Data Pipeline for End-to-End Mathematical Reasoning
Abstract
Training a strong mathematical reasoning model from a base model requires three stages—mid-training, SFT, and RL—with fundamentally different data needs. Most prior work curates data one stage at a time, so locally optimal choices fail to compose end-to-end; we instead take a joint multi-stage view, placing every problem where it is most valuable to the final model. We instantiate this view with daVinci-Math, a unified pipeline built on two pillars: (i) single-source preparation—cleaning, deduplication, and filtering shared across all three stages are applied once to a corpus aggregated from public sources, producing a 2.7M-problem pool; and (ii) stage-aware routing—each problem is labeled via multi-criteria classification and answer verification, then routed to mid-training, post-training, or filtered out, with post-training further split into SFT and a rule-verifiable RL subset. The pipeline yields 62B mid-training tokens, 3.8M SFT trajectories, and 39K RL prompts. We instantiate the full pipeline at both the 3B and 30B-A3B scales: daVinci-Math-3B reaches 86.6 on AIME 2025 and 75.1 on HMMT Feb 2025, and daVinci-Math-30B-A3B reaches 90.4 average over four benchmarks, setting a new state of the art at each scale. Controlled head-to-head comparisons further show that daVinci-Math provides stronger training data than open math datasets at both the mid-training and SFT stages, and ablation studies confirm the effectiveness of our pipeline design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.