Learning in Language and Weights: When the Two Reinforce Each Other
Abstract
Language models are unique in that they can learn both from context and weight optimization, each with complementary strengths. Here, we study a method for combining them: the sample efficiency and interpretability of learning from natural language, and the permanence of weights learning. We show that CORE, a nonparametric learning algorithm, can produce natural language insights that can be directly distilled back into weights via insight distillation. Across six reasoning tasks and two models, we find that our method is more sample efficient and outperforms rejection-sampled fine-tuning on all 12 pairs, and tuned GRPO baselines on 10/12 pairs, and uses the distilled insights while solving problems. A further round of CORE learning improves on the previous round of insight distillation in 11/12 pairs, demonstrating a mutually reinforcing relationship between language and weights learning. Insight distillation is more interpretable than in-weights learning alone, and allows direct control on what a model learns: we are able to easily detect a reward hacking strategy occurring in the wild in a coding task, and remove the exploiting insights before distillation to prevent the behavior from forming. Finally, we show signs of life with an autonomous loop that can schedule repeated rounds of insight distillation for further learning efficiency. Overall, our results show that with powerful enough language models, learning in natural language and with weights can reinforce each other.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.