EMBEDDING-ONLY KNOWLEDGE TRACING: PREDICTING RESPONSES WITHOUT KNOWLEDGE COMPONENT ANNOTATIONS
Abstract
Knowledge tracing (KT) predicts a learner's next response from the response history, and adaptive systems use it to choose the next question and to find learners who need support. Standard KT methods rely on expert-defined knowledge components (KCs) with a costly concept vocabulary and on a learned question-ID table, which has no row for questions absent from training. Embedding-only Knowledge Tracing (EoKT) feeds fixed general-purpose question-text embeddings to a compact recurrent predictor, so it needs neither KC annotations nor an ID table. On MOOCRadar and XES3G5M, we change only the question-side input of this predictor and score warm questions and cold questions withheld from training. In standard AUC, which pools test responses to almost entirely seen questions, EoKT reached the validation-selected KC-based method, with paired differences of +0.0056 on MOOCRadar and +0.0046 on XES3G5M under unequal search budgets. In the same predictor, a learned ID table was at least as good on warm questions, and adding learned or pretrained KC inputs did not raise AUC. In cold AUC, which ranks each learner's responses to questions withheld from training, text embeddings reached 0.6041 on MOOCRadar and 0.6489 on XES3G5M, against 0.5047 and 0.5896 for learned IDs, which cannot distinguish withheld questions; the ordering held for three embedding models. Text embeddings from gemini-embedding-2 also exceeded permuted embeddings on both datasets by +0.0848 and +0.0833 in cold AUC, with 95% intervals that exclude zero when learners and questions are both resampled, so matching each question to its own text contributes to the gain. With question representations precomputed, EoKT returns one prediction from a retained state in about 0.50 ms on a single CPU thread. EoKT can be trained across datasets by simple concatenation: one model pooled over three datasets that differ in language and subject lost at most 0.0012 AUC against per-dataset models in a single-fold check. These results support EoKT as a practical KC-free knowledge tracing method whose text input also supports prediction on questions absent from training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.