MILE: Milestone-Guided Reinforcement Learning for Mathematical Reasoning
Abstract
Reinforcement learning has shown strong potential for improving the mathematical reasoning capabilities of large language models (LLMs). However, existing methods often depend on sparse outcome feedback, manually designed process signals, or separately trained reward models, leaving expert reasoning trajectories underexploited as a rich source of dense supervision. We propose MILE, a milestone-guided reinforcement learning framework that models reasoning progress through intermediate states characterized by their future solvability. MILE extracts milestones from multiple expert solutions and constructs dense rewards for student reasoning transitions by measuring both progress and directional consistency. These rewards are incorporated into a two-stage GRPO optimization framework, enabling fine-grained transition-level learning followed by trajectory-level reasoning optimization. Experiments on mathematical reasoning benchmarks show that MILE consistently outperforms representative RL and distillation baselines. Further analyses validate the effectiveness of its components, scalability with expert trajectories, and improved reasoning efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.