OmniEduBench: A Chinese Dataset for Evaluating Subject Knowledge and Pedagogical Judgment
Abstract
Evaluating large language models (LLMs) for educational use requires distinguishing subject knowledge from judgment about how to respond to learners. Knowledge-oriented assessments provide limited evidence about such pedagogical judgment, motivating complementary evaluation in Chinese educational contexts. We introduce OmniEduBench, a Chinese dataset for evaluating subject-knowledge problem solving and scenario-based pedagogical judgment. It comprises approximately 24.6K question–answer pairs spanning 41 academic subjects and 20 pedagogical competencies. The knowledge component covers 11 question formats across multiple educational stages, while the pedagogical component primarily uses multiple-choice scenarios involving instructional guidance, feedback, and learner support. The dataset combines public educational resources, private institutional materials, and model-generated scenarios, followed by cleaning, model-based selection, and human review. We evaluate 15 LLMs and analyze the sensitivity of open-ended answer scoring to the choice of LLM judge. Results show substantial variation across subjects and pedagogical categories, while evaluator choice materially affects absolute scores. Findings highlight the need for scoring criteria and evaluator validation when interpreting educational benchmarks. OmniEduBench provides a resource for diagnosing subject-knowledge and pedagogical-judgment limitations, with pedagogical scores interpreted as scenario-based judgments rather than direct measures of teaching effectiveness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.