BanduraBench: The Mechanism Mirage in Theory-Grounded Educational LLM Reasoning
Abstract
Large language models are often evaluated by whether they produce the correct final answer, but a correct decision does not necessarily imply that the model identified the right evidence or recovered the mechanism supporting it. We introduce BanduraBench, a 4,530-item benchmark grounded in Bandura's Social Cognitive Theory of Self-Efficacy for evaluating the gap between decision accuracy and mechanism recovery in educational reasoning. BanduraBench combines foundational construct evaluation with contextualized instructor- and researcher-facing reasoning tasks constructed from an expert-reviewed ontology, human-verified seeds, gold-first blueprints, and human-adjudicated admissible pathways. Models are evaluated on explicit, externally specified mechanism components—Bandura source, task-specific online-learning self-efficacy, course-design interpretation, and causal-claim strength. Across 15 LLMs, DesignReasoning answer accuracy is nearly saturated at 96.6% on average (95.3-98.2%), while normalized four-component recovery averages only 65.5% and ranges from 45.1-84.7%. Strict four-component recovery averages 21.4%, showing that much of the strict-score reduction follows from conjunctive accumulation of component errors. BanduraBench therefore provides a conjunction-aware framework for evaluating whether similar educational decisions are supported by similarly coherent theory-grounded representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.