acceptodds
Under review as a conference paper at ICLR 2027

MILE: Milestone-Guided Reinforcement Learning for Mathematical Reasoning

Abstract

Reinforcement learning has shown strong potential for improving the mathematical reasoning capabilities of large language models (LLMs). However, existing methods often depend on sparse outcome feedback, manually designed process signals, or separately trained reward models, leaving expert reasoning trajectories underexploited as a rich source of dense supervision. We propose MILE, a milestone-guided reinforcement learning framework that models reasoning progress through intermediate states characterized by their future solvability. MILE extracts milestones from multiple expert solutions and constructs dense rewards for student reasoning transitions by measuring both progress and directional consistency. These rewards are incorporated into a two-stage GRPO optimization framework, enabling fine-grained transition-level learning followed by trajectory-level reasoning optimization. Experiments on mathematical reasoning benchmarks show that MILE consistently outperforms representative RL and distillation baselines. Further analyses validate the effectiveness of its components, scalability with expert trajectories, and improved reasoning efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.