StudentBench: AI and human tutoring yield equivalent GRE learning gains
Abstract
AI offers an unprecedented opportunity to augment human capabilities by teaching us new knowledge and skills. However, the extent to which AI can do this had not been measured at scale. We introduce StudentBench, a suite of LLM teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student–AI messages to study whether LLMs produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that pooled AI tutoring is statistically equivalent to expert human tutoring for combined GRE learning gains (). In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. The two studies clearly separate LLMs for tutoring across (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. One model achieved learning gains equivalent to human tutoring () at lower cost per percentage point of learning gain. For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.