The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning
Abstract
Large language models (LLMs) trained to answer questions are natively poor at teaching, and Reinforcement Learning (RL) training with a simulated student is a promising approach to improve the pedagogical quality of LLMs. Existing RL-trained tutors reward the student's success on the tutored problem with the tutor's words still in context; thus, the reward is easiest to raise by telling students answers, and a tuned penalty is needed to reduce telling. This work studies this simulation-based training approach for LLM tutors and how the learning sciences-inspired approach can let LLMs learn to teach more effectively. We introduce a masked near-transfer post-test: the student is tested on an unseen variant of the tutored problem with the tutor's utterances masked, so the reward can rise only through the student's own turns. This prevents cognitive offloading by the student and allows the continuous penalty to be replaced by two binary reward gates (factual correctness of tutor response, no solution handover). A leave-one-out ablation shows that the learning-gain reward on its own does not separate teaching from telling, while binary reward gates suppress solution handover, the near-transfer post-test improves out-of-domain transfer, and removing tutor utterance masking coincides with a rise in thinking tokens. Using these reward designs, we develop Eduardo, a recipe for training LLM tutors, and use it to train 4B, 9B, 14B and 27B models from two distinct LLM architectures. Our post-trained Eduardo-27B model matches Gemini-3.1-Pro on MathTutorBench and Claude-Opus-4.8 on TutorMoments at 2.4-6.2x fewer thinking tokens (a requirement for interactive tutoring), demonstrates an increase in push for justification teacher moves without it being named in the reward, and trains out support fading (e.g., assigning independent work), whose payoff lies beyond a single-problem dialog episode. We open-source our training environment, an 8,671-problem near-transfer dataset, and trained models for further development.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.