StudentSim: Training LLM-based Student Simulators
Abstract
AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but feedback about which guidance works for whom is sparse, slow, and costly to collect. Student simulators can provide this feedback at machine timescales. Existing approaches capture only part of what tutor training requires: state-tracking models fit learner behavior but cannot respond to natural-language guidance, while prompted LLMs respond fluently but do not reliably reproduce an individual student's competence and characteristic errors. We present StudentSim, a training framework that turns sparse per-student records into individualized simulators through pooled domain training followed by per-student specialization. StudentSim combines structured behavioral grounding with a language-model interface, enabling each simulator to mirror a student's responses and revise them under tutor guidance. We also introduce StudentSimEval, a standardized protocol over 60 students in chess, second-language English writing, and mathematics. It evaluates behavioral fidelity (F↑), agreement between the simulator's responses and the real student's responses, and guidance responsiveness (R↑), whether the simulator reaches the response targeted by guidance. Across all three domains, StudentSim outperforms behavior-prediction and prompted-LLM baselines on both metrics. In chess, StudentSim reaches F = 0.51 and R = 0.91, compared with 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. To demonstrate downstream use for AI tutor improvement, we use a trained StudentSim as the reward in tutor reinforcement learning. Expert chess players rate the resulting tutor as more accurate, better-guided, and more personalized than a no-RL baseline and a tutor trained with GPT-5.4 simulator feedback.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.