Investigating the Robustness of Knowledge Tracing Models in the Presence of Shifting Student Distributions
Abstract
Knowledge Tracing (KT) has been an established problem in the educational data mining field for decades, and it is commonly assumed that the underlying learning process being modeled remains static. Given the ever-changing landscape of online learning platforms (OLPs), we investigate how concept drift and changing student populations can impact student behavior within an OLP through testing model performance both within a single academic year and across multiple academic years. Four well-studied KT models were applied to five academic years of data to assess how susceptible KT models are to concept drift. Through our analysis, we find that all four families of KT models can exhibit degraded performance. Bayesian Knowledge Tracing (BKT) remains the most stable KT model when applied to newer data, while more complex, attention-based models lose predictive power significantly faster when features provided are diverse enough to allow overfitting. Finally, we apply well-studied measures of dataset similarity to assess which kinds of drifts may account for the performance degradation we document. To foster more longitudinal evaluations of KT models, the data used to conduct our analysis is available at [this link](https://osf.io/hvfn9/?view_only=b936c63dfdae4b0b987a2f0d4038f72a)
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.