acceptodds
Under review as a conference paper at ICLR 2027

ScaffoldBench: A Dynamic Benchmark for Evaluating Pedagogical Scaffolding and ZPD Alignment in Large Language Models

Abstract

A language model can solve a problem while failing to help a student solve it. This mismatch exposes two distinct requirements for tutoring: inferring the learner's cognitive state and generating assistance that preserves the learner's remaining reasoning. We introduce ScaffoldBench, a benchmark that separates cognitive objective recognition, structured state diagnosis, and constrained instructional generation. It comprises 2,000 expert-verified synthetic student traces built from physics and reading problems, with explicit knowledge chains, cognitive state vectors, and error categories. We formalize the relationship between local and global diagnostic accuracy, derive the leakage-imposed ceiling on scaffolding fidelity, and distinguish observed Oracle gains from the isolated value of diagnostic information. Evaluation of 19 models reveals a diagnosis–intervention mismatch: high diagnostic accuracy can persist under mismatched student traces, while reasoning-oriented models exhibit substantial answer leakage even with explicit state information. On physics tasks, their mean leakage rate is 91.0% under baseline prompting and 78.9% under the augmented Oracle protocol. Educational models respond much more strongly to this protocol, reaching 1.0% leakage and 4.89/5 scaffolding fidelity on reading tasks. These results establish cognitive diagnosis, trace dependence, and response restraint as separate axes of tutor evaluation, motivating models that infer learner states and control what they reveal.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.