Co-evolving Tutor and Student Simulator for Dialogue Tutoring
Abstract
Optimizing large language models for dialogue tutoring requires feedback about how guidance changes students' reasoning and problem-solving progress. An objective inspired by the zone of proximal development (ZPD) favors guidance associated with intermediate predicted response correctness. Our analysis identifies the largest simulation errors at intermediate reference correctness. As teaching policies evolve, student models must continue to represent the mistakes and incomplete reasoning that instructional guidance is intended to address. We propose COMET, an online reinforcement-learning framework that couples instructional utility optimization with student-response modeling through policy-induced dialogue histories. Simulator validation performance, in turn, controls the granularity of the Tutor's ZPD reward. The simulator objective combines cognitive rewards with reference-based behavioral constraints, jointly targeting reasoning quality and agreement with recorded student responses. A shared judge-anchored evaluator refines its rubric and distills rubric-based judgments into a prediction head for efficient cognitive scoring. COMET achieves 85.3% Clean Success@15. With an independent frozen student and a shared answer-disclosure guard, it improves Success@15 over a same-base prompted tutor by 20.1 percentage points. These findings demonstrate the effectiveness of tutor–student co-evolution for learning tutoring policies through self-play, with instructional gains that transfer to an independent student simulator.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.