Who Scores the Tutor? Sampling Noise and Self-Preference in Simulated-Learner Evaluation
Abstract
Tutoring agents are compared by the learning gain they produce in a simulated learner. This learner is a knowledge-tracing model fitted to real interaction data, and the same model also scores the teaching. On one corpus, both DKT scorers rank a policy that plans against a DKT model first out of eight, but an AKT scorer ranks it seventh on the same learners and sequences. This paper isolates the effect of the scorer. Five planning policies (one per knowledge-tracing architecture) and three scripted controls choose exercises for the same learners. The chosen sequences are then frozen, and 10 scorers (five architectures with two seeds each) rank the policies on each of three corpora. Disagreement between scorers is measured against a resampling floor, the agreement of one scorer with itself on disjoint halves of the learners. This floor is already low (Kendall ), so with sampling noise alone a benchmark of this size barely determines its own ranking. Changing the architecture of the scorer (scorer identity) lowers by a further (bootstrap 95% interval , 2,062 simulated learners from 999 students). On the two ASSISTments corpora, changing only its seed adds no disagreement beyond resampling. Scorers also show a directional bias. A scorer ranks the policy planned against its own architecture positions higher than scorers of other architectures do (exact assignment tests, ). The bias is positive on all three corpora and depends on the planning. A greedy control that uses the same model without planning against it receives no detectable preference. On ASSIST2017, raising the search budget from 12 to 192 candidates nearly triples the gain that scorers of the same architecture as the planner see in excess of the other scorers. Averaging a larger panel of scorers does not produce a usable ranking on average at any panel size tested. Excluding scorers of the same architecture as the planner is free and reorders every corpus leaderboard. It changes the reported winner on NIPS34 in both runs and on ASSIST2017 in the second run. The protocol, exclusion estimator and instrument will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.