ReTraceKT : Relevance-guided Multimodel State-space Knowledge Tracing
Abstract
Knowledge tracking (KT) estimates a learner's evolving knowledge state from a sequence of exercise responses. Two properties of modern learning platforms make this problem harder: essential evidence is often embedded in diagrams rather than text, and long histories contain many interactions that are irrelevant to the next exercise. We present ReTraceKT, a multimodal KT framework that retrieves semantically relevant evidence before tracking knowledge dynamics. ReTraceKT first uses a vision-language model offline to render each exercise image into structured descriptions and aligns them with trainable exercise embeddings. It then builds a threshold semantic graph over exercises and uses the graph to retrieve a sparse, question-conditioned summary of a learner's history. Finally, a selective state-space tracer combines local response patterns with long-range state transitions in linear sequence complexity. Experiments on NIPS34, MOOC-Radar, and XES3G5M show consistent improvements over strong recurrent, attention-based, and sparse KT baselines. Relative to the strongest baseline, ReTraceKT improves AUC by 3.22%, 1.82%, and 2.08% on the three datasets, respectively; ablations confirm that semantic rendering and relevance retrieval are complementary. Overall, ReTraceKT delivers consistently strong and stable performance across the three datasets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.