Can the Simulated Student Learn? Fitted Knowledge-Tracing Models Disagree on Spacing, Recency and Transfer
Abstract
Tutoring agents are increasingly scored against a simulated student. In benchmarks that track what the learner knows, the simulated student is a knowledge-tracing model fitted to interaction logs. The score is the change in the predicted probability of a correct response after teaching. It measures learning gain only if the model responds to teaching as a learner does. This paper introduces four contrasts that test this condition on five architectures fitted to two corpora at three seeds each. The contrasts are spacing against a position-matched control, recency, transfer from a related concept and the error penalty. Each contrast appends two sequences to the same learner history. They differ in only one systematic respect: the order of practice, the concept of the items practised first or whether the first response is correct. The effect is the paired difference in predicted gain on held-out items of the target concept, averaged across histories. A contrast passes when its effect has the expected sign at in a two-sided paired -test. That sign is positive for the three schedule contrasts and negative for the error penalty. A failure is an absence if the effect is not significant and an inversion if it is significant the other way. Responses are written in rather than sampled, because they move the probe score a median of 10 times more than the concept practised does. Every simulator penalises a wrong answer at every seed. Spacing and recency each fail in six of 10 simulators at the first seed, and transfer fails in eight. Across the three seeds, some absences and one small inversion persist, but the largest inversions vanish. Only five of 10 simulators give the same verdicts at every seed, and new target concepts change verdicts at least as often as a refit. DKVMN fails spacing at every seed and target concept. Validating DKT on ASSIST2017 alone finds no failure at the first seed. When the simulator answers, as in a benchmark, a spaced tutor beats a massed tutor on 16 of 30 fitted models, ties with it on 12 and loses on two. The fitted simulator that a benchmark adopts therefore limits which tested phenomena the benchmark can credit. A four-contrast certification costs minutes and reports which of the four a fitted simulator registers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.