Beyond Problem-Solving: A Multi-Agent Framework and Pedagogically-Aligned Benchmark for Evaluating LLMs’ Educational Question Generation
Abstract
Large Language Models (LLMs) have achieved impressive performance in educational applications, particularly in solving mathematical problems. However, existing evaluation benchmarks overwhelmingly emphasize solution correctness while neglecting the equally important capability of generating high-quality educational questions. This results in a fundamental gap in assessing LLMs’ pedagogical competence and prevents a complete teaching–learning evaluation loop. We address this gap by introducing a multi-agent mathematics Question Generation (QG) framework that explicitly decomposes teacher-like behaviors—such as intent interpretation, curriculum planning, difficulty calibration, and solution verification—into collaborative LLM agents, and the first pedagogically aligned, multi-dimensional benchmark specifically designed to evaluate question generation quality. The benchmark assesses generated questions along five major dimensions—Accuracy, Difficulty, Novelty, Educational Value, and Human Preference Alignment—covering 19 fine-grained indicators grounded in educational psychology and assessment theory. Experimental results show that our multi-agent framework consistently outperforms strong LLM baselines across all five dimensions, with particularly large gains in higher-order pedagogical metrics such as diagnostic value, reasoning depth, and instructional clarity. Overall, this work takes an important step toward systematic and pedagogically grounded evaluation of LLM-driven educational content generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.