acceptodds
Under review as a conference paper at ICLR 2027

TutorSim: Evaluating the Tutoring Capabilities of Large Language Models

Abstract

We propose a novel simulation environment for evaluating the tutoring capabilities of large language models (LLMs). Our simulation is based on data from a randomized controlled trial testing two different LLM tutors on several hundred high school students. Based on this data, we design two key components: (1) an LLM prompted to mimic student queries, and (2) an LLM rubric calibrated to estimate the grade a student achieved based on their conversation with the tutor. Using our evaluation, we perform two empirical analyses on existing LLM tutors. First, we ablate different portions of a system prompt focused on tutoring. We find that portions of the prompt focusing on avoiding giving away answers are most critical for achieving good performance; surprisingly, portions suggesting tutoring strategies actually reduces performance, suggesting that the LLM is more effective at identifying good tutoring strategies on its own. Second, we ablate the size of the underlying LLM. Surprisingly, we find that the medium-sized LLM is a more effective tutor, suggesting that LLM errors might actually improve learning by forcing students to think. We also find that the ranking is inconsistent with the na\"ive LLM judge for assessing tutoring capabilities, highlighting the need to calibrate tutoring performance estimates to rigorous outcomes such as exam grades. Our simulation-based results complement rigorous causal evaluations on LLM-based tutors by enabling a more systematic evaluation of the impact of different design decisions on the tutoring capabilities of LLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.