EduMCTS: Improving Large Language Model Reasoning through Adaptive Pedagogical Interaction
Abstract
High-quality reasoning trajectories serve as critical post-training resources for enhancing the reasoning abilities of large language models (LLMs). However, existing reasoning data augmentation methods mainly focus on exploring diverse candidate solutions and selecting among reasoning trajectories, while overlooking how to actively improve intermediate reasoning during trajectory generation. To mitigate this limitation, we propose EduMCTS, a reasoning data augmentation framework that constructs reasoning trajectories by selectively introducing pedagogical interventions during problem solving. Specifically, EduMCTS formalizes pedagogical strategies as executable search actions for intermediate reasoning states, evaluates candidate reasoning paths using local reasoning progress and final answer correctness, and leverages tree search to construct verified interactive reasoning trajectories. Extensive experiments across three LLMs and eleven reasoning benchmarks, spanning seven mathematical and four scientific reasoning tasks, demonstrate that EduMCTS consistently outperforms existing data augmentation methods and generalizes well to out-of-domain scientific reasoning. Further analyses support the role of pedagogical strategy search and show that models trained with EduMCTS exhibit adaptive pedagogical reasoning behaviors during inference. Our code and dataset are anonymously available at https://anonymous.4open.science/r/EduMCTS-B568/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.