Exam2Teach: Can LLM Agents Turn Exam Results into Teaching Materials?
Abstract
High-quality teaching materials empower teachers to bridge learning gaps and provide highly targeted student support. Preparing such materials from exam results requires accurately interpreting student answers, translating identified needs into targeted instruction, and producing clear, coherent multimodal documents. In this work, we introduce Exam2Teach, a benchmark for evaluating whether autonomous LLM agents can generate high-quality, personalized teaching materials. Exam2Teach spans five core subjects shared by university-entrance curricula worldwide, using real exam papers, grading rubrics, de-identified student answers and scores. We design an open-source multi-agent framework for composing multi-step agent workflows. It powers our entire benchmark: agents generate documents, and evaluation workflows combine rule-based checks with agentic inspection of source evidence and rendered documents to assess task completion, layout quality, factual correctness, and personalization. The framework is general-purpose and extends readily to new tasks. An extensive evaluation of 10 agent configurations across 150 runs reveals uneven performance across four dimensions: the task completion threshold was met in 45.3% of runs, layout quality in 62.0%, factual correctness in 12.7%, and personalization in 26.7%. Only 7.3% passed all four dimensions, with no passing runs for review slides, and no single model consistently produced high-quality teaching materials of all categories. Even the strongest model overall, GPT-5.5, passed all four dimensions in only 40.0% of its runs. These results indicate a clear gap for frontier LLM agents in teaching material generation for AI-assisted education. Code and data are available at https://anonymous.4open.science/r/Exam2TeachAnonymous-8E75/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.