Evaluating the Ability of Large Language Models to Generate Structured Curricula for Human Learners
Abstract
Large language models (LLMs) are increasingly used to teach human learners, but it remains unclear what aspects of their instruction support learning. Most work has focused on interactive tutoring, where models can respond to questions, provide feedback, and adapt to a learner. Here, we isolate a complementary component: the curriculum a model chooses to present in advance. We evaluate LLM-generated curricula with 1,428 human learners across four concept-learning settings, progressing from controlled tasks in which teaching strategies can be characterized formally to more complex scientific and mathematical subjects. Where tractable, we compare LLM curricula with Bayesian pedagogy, a normative account in which teaching depends on selecting information with a learner's inference in mind. LLM-generated curricula align with the Bayes-optimal model, while exhibiting systematic preferences among distinct strategies that the theory treats as equally optimal, and these preferences benefit human learners. As the concepts become more compositional, models spontaneously introduce intermediate concepts and organize them into structured curricula that improve learning. The same organization appears in naturalistic subjects, where static LLM-generated curricula substantially improve human performance. Together, these results suggest that an important component of LLM teaching lies in how models select and structure instructional content, and that this organization extends from choosing examples to decomposing complex concepts into learnable parts for human learners.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.