acceptodds
Under review as a conference paper at ICLR 2027

Context into Weights: Self-Distillation for Teaching a Small Model a Programming Language It Cannot Write

Abstract

Small language models can transfer algorithmic knowledge across programming languages, yet may fail to express that knowledge in languages whose syntax and libraries are poorly represented in their pretraining. The standard post-training recipe, supervised fine-tuning (SFT) followed by reinforcement learning with group relative policy optimisation (GRPO), breaks at both stages here: there is no high quality corpus to fine-tune on, and when no sampled program passes its tests, GRPO has no learning signal. Yet the missing knowledge is small and usable in context: with a short syntax sheet in its prompt, base Qwen3-8B solves 28 of 30 elementary Ballerina programming language problems, and none without it. We therefore replace both stages with self-distillation, in which the model with the language documentation in its context teaches the same model without it. Self-distilled supervised fine-tuning (SD-SFT) fine-tunes the model on its own documentation-conditioned programs that pass the tests. Self-distillation policy optimisation (\sdpo) then trains every token of the model's own attempts toward a teacher that also sees the documentation and the compiler's feedback. On Ballerina, where Gemma 4 E4B scores 0% pass@1, \sdsft reaches 55.4% and \sdpo 62.7%. The same incremental ordering holds across other models and languages. Because the recipe needs no corpus in the target language, only a compiler, tests and the syntax sheet, it offers a route to adapting small models to many languages that pretraining leaves out. We release all the language agnostic training data, implementation and the model weights.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.