acceptodds
Under review as a conference paper at ICLR 2027

Context Distillation Distills More Than Skills: Navigating Skill Gains and Cross-Domain Retention with State-Shared Drift

Abstract

Skill documents provide language model agents with reusable procedural knowledge, but retrieving and processing them at inference time incurs computational cost. On-policy context distillation can internalize such knowledge by training a document-free student to match a document-conditioned teacher. However, this process can improve the target skill while degrading instruction and rule following in unrelated domains. We identify state-shared drift: systematic token-score shifts that recur across prediction contexts with context-dependent magnitudes. Under heterogeneous distillation, we decompose this drift into document-induced and model-induced sources whose directions and scales differ substantially. Larger state-shared drift coincides with greater cross-domain degradation, while progressively attenuating it traces a trade-off between document-free skill gains and cross-domain retention. Based on this, we propose Drift-Corrected Skill Distillation (DCSD), which separately estimates the two source-specific drift components and uses them as correction directions. Control documents and selective masking help preserve skill guidance, while correction strength and correcting one or both sources adjust the balance between skill gains and cross-domain retention. Experiments with Qwen3 models on SRA-Bench and six cross-domain instruction- and rule-following benchmarks show that DCSD substantially improves cross-domain retention while preserving useful document-free skill gains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.