TutorSim: Evaluating the Tutoring Capabilities of Large Language Models
Abstract
We propose a novel simulation environment for evaluating the tutoring capabilities of large language models (LLMs). Our simulation is based on data from a randomized controlled trial testing two different LLM tutors on several hundred high school students. Based on this data, we design two key components: (1) an LLM prompted to mimic student queries, and (2) an LLM rubric calibrated to estimate the grade a student achieved based on their conversation with the tutor. Using our evaluation, we perform two empirical analyses on existing LLM tutors. First, we ablate different portions of a system prompt focused on tutoring. We find that portions of the prompt focusing on avoiding giving away answers are most critical for achieving good performance; surprisingly, portions suggesting tutoring strategies actually reduces performance, suggesting that the LLM is more effective at identifying good tutoring strategies on its own. Second, we ablate the size of the underlying LLM. Surprisingly, we find that the medium-sized LLM is a more effective tutor, suggesting that LLM errors might actually improve learning by forcing students to think. We also find that the ranking is inconsistent with the na\"ive LLM judge for assessing tutoring capabilities, highlighting the need to calibrate tutoring performance estimates to rigorous outcomes such as exam grades. Our simulation-based results complement rigorous causal evaluations on LLM-based tutors by enabling a more systematic evaluation of the impact of different design decisions on the tutoring capabilities of LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.