CogTeachBench: A Standardized Learner-State Benchmark for Evaluating LLM Tutors
Abstract
Large language models (LLMs) are increasingly embedded in learners’ study processes as explainers, feedback providers, and personalized tutors. However, evaluating whether an LLM can actually teach remains challenging, especially when teaching is treated as measurable change in a learner’s misconception state. Existing educational benchmarks typically assess static response quality, task-solving accuracy, or isolated pedagogical rubrics, while learner variation is often reduced to coarse role prompts or ability labels. To address this gap, we introduce CogTutorBench, a benchmark for evaluating LLM tutors through controlled misconception repair. We instantiate it in algorithmic programming using LiveCodeBench problems released after the learner-proxy models’ pretraining cutoff, resulting in 316 oracle-verified misconception-repair scenarios. We further define standardized learner profiles by varying ability tier, acceptance behavior, and cognitive-load constraints, and evaluate tutors with four complementary metrics, including Cognitive Bottleneck Clearance Score (CBCS), Adaptive ZPD Span (AZS), Process-level Scaffolding Quality (PSQ), and Pedagogical Cognitive Efficiency (PCE). Experiments show that tutoring quality has multiple rankings. Gemini 3.1 Pro achieves the strongest post-teaching clearance, while GPT-5.4 achieves the highest process quality and load-normalized efficiency. Learner profiles substantially change measured outcomes. Cognitive-load constraints reduce teaching gain, whereas high-acceptance or higher-ability learners may produce fluent agreement signals that some tutors mistake for executable understanding. These findings suggest that robust LLM tutor evaluation must jointly measure learner improvement, pedagogical process, cognitive efficiency, and variation across learner profiles.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.