acceptodds
Under review as a conference paper at ICLR 2027

Co-evolving Tutor and Student Simulator for Dialogue Tutoring

Abstract

Optimizing large language models for dialogue tutoring requires feedback about how guidance changes students' reasoning and problem-solving progress. An objective inspired by the zone of proximal development (ZPD) favors guidance associated with intermediate predicted response correctness. Our analysis identifies the largest simulation errors at intermediate reference correctness. As teaching policies evolve, student models must continue to represent the mistakes and incomplete reasoning that instructional guidance is intended to address. We propose COMET, an online reinforcement-learning framework that couples instructional utility optimization with student-response modeling through policy-induced dialogue histories. Simulator validation performance, in turn, controls the granularity of the Tutor's ZPD reward. The simulator objective combines cognitive rewards with reference-based behavioral constraints, jointly targeting reasoning quality and agreement with recorded student responses. A shared judge-anchored evaluator refines its rubric and distills rubric-based judgments into a prediction head for efficient cognitive scoring. COMET achieves 85.3% Clean Success@15. With an independent frozen student and a shared answer-disclosure guard, it improves Success@15 over a same-base prompted tutor by 20.1 percentage points. These findings demonstrate the effectiveness of tutor–student co-evolution for learning tutoring policies through self-play, with instructional gains that transfer to an independent student simulator.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.