acceptodds
Under review as a conference paper at ICLR 2027

Contrastive Co-Knowledge: Dynamic Learner Embeddings for Vocabulary Knowledge Prediction

Abstract

Vocabulary inventory prediction asks whether a learner knows a target word, given only a small set of probed words. Existing approaches model this as scalar latent ability (item response theory) or as classification over hand-engineered word features (frequency, age of acquisition, concreteness). Both treat word representations as fixed, inheriting a metric optimised for distributional semantics rather than acquisition. We argue that the missing ingredient is the metric itself. We introduce (Contrastive Co-Knowledge Embeddings), which uses the words a learner knows together as weak supervision for adapting a pretrained text encoder: LoRA-based contrastive fine-tuning on co-known pairs pulls together words that are acquired together, with frequency-matched and learner-specific hard negatives that prevent the objective from degenerating into a frequency predictor. A learner is then a point in this space—an aggregation of their known-word embeddings—paired with a signed proficiency coordinate: we prove that no non-negative pooling, in any dimension, can place a learner far from the whole lexicon, while a single extra coordinate spans zero knowledge to near-mastery. Prediction is a two-tower score decomposing into word-intrinsic difficulty, personalised interaction, and proficiency offset. Because the learner point is a running aggregate, it admits exact online updates and a principled cold-start shrinkage rule. On two public “tall” benchmarks (, ; , ) and a large simulated benchmark, improves per-learner macro-AUC by 3.9–4.2 points over the strongest non-embedding baseline and by 6.3 points over the same encoder used without adaptation, with gains concentrated on low-frequency words where word-prior models saturate. all numbers illustrative We release the full protocol—probe sampling, leave-one-learner-out splits, leakage controls—so that the small-learner, many-word setting can be benchmarked reproducibly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.