acceptodds
Under review as a conference paper at ICLR 2027

Predictive Anchoring Shapes Contextual Dependence in Language-Model Pretraining

Abstract

We use knowledge distillation to study learned contextual dependence during language-model pretraining by varying teacher context regimes and composition, while holding the student's architecture and input window fixed. A frozen short-context teacher supplies a pressure to match locally conditioned predictions, alongside ordinary next-token learning. This short-context distillation (sKD) improves baseline predictions across WikiText-103 and SlimPajama-100M, while attenuating average contextual effects both where distant text helps baseline prediction and where it hurts. An equal-context distillation (eqKD) control, whose teacher receives the same causal input as the student, achieves lower overall loss than sKD but retains more of both helpful and harmful influences. We observe a tradeoff between helpful-context use and harmful-context attenuation in these contrasting distillations. Mixed-teacher distillation (mixKD) trains a single student using a fixed mixture of the two teacher distributions. mixKD achieves overall prediction loss close to eqKD, with lower loss in most, but not all, comparisons. It demonstrates substantially improved prediction on helpful positions compared with sKD, while retaining a prediction advantage over eqKD on harmful positions. It also shows greater average contextual benefit on helpful positions than sKD and a smaller average contextual penalty on harmful positions than eqKD. These findings hold across three seeds on validation and test data, at two evaluated short-context horizons on WikiText-103 and one on SlimPajama-100M. Equal-weight ensembles of separately trained sKD and eqKD students achieve lower overall prediction loss than mixKD, yet an exploratory post-test analysis finds that no fixed probability weighting of these student pairs simultaneously matches or improves on both mixKD's helpful-context benefit and harmful-context penalty. This study shows that similar overall prediction loss can accompany different learned contextual behavior, which can be shaped by teacher context regimes and composition.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.