acceptodds
Under review as a conference paper at ICLR 2027

Optimizing Tutor LLMs for Pedagogy While Maintaining Student Engagement

Abstract

Large language models for education are optimized to follow effective pedagogy: they scaffold student learning and help students evaluate and self-correct their reasoning, without giving direct answers. These tutors are typically trained with supervised learning and reinforcement learning against a simulated student that is an assistant LLM. Such simulated students are always enthusiastic and respond at every turn, unlike real students. As a result, when a student becomes less cooperative, current tutors give away solutions too often and struggle to keep the learner engaged. To address this, we train tutors with reinforcement learning against a user model: an LLM trained to imitate human users, rather than being a helpful assistant. Like real students, it can lose interest and quit the conversation at any time, so the tutor must keep it engaged to teach it anything. We train using a combination of two rewards: an engagement reward to keep the student's attention, and a learning reward to maintain pedagogy. We find that the two rewards are complementary, because staying engaged enables longer conversations where learning happens in later turns. Overall, our tutors achieve higher judged learning than prior tutor models while leaking solutions less. We will release two open-source tutors, one for easier and one for harder math problems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.