Verify In-Sample, Withhold, Distill: Three Insights for LLM Skill Learning
Abstract
Skill learning has become a popular approach to improving LLM agents: for each training instance, a model reflects on the policy's prediction against the ground truth and writes one natural-language lesson; these lessons are then consolidated into reusable skills that the policy consults at inference time. Some recent work draws a productive analogy with backpropagation. But the two also differ in fundamental ways, and these differences give rise to three insights. 1) Verify in-sample}: a lesson is worth keeping only if it fixes its own training sample. It is common to verify a lesson on a validation set, but that is an out-of-sample test, measuring generalization rather than whether the lesson repairs the sample it was derived from. Verifying in-sample is critical: unlike a sufficiently small gradient step, which by construction reduces the loss on the sample it is computed from, a lesson offers no such guarantee. A lesson should therefore be iteratively fed back to the policy and revised until it fixes its own training sample, and discarded otherwise. 2) Skill learning overfits far less than backpropagation. A gradient step fits whatever reduces the loss; an individual lesson can overfit its own sample too, but this overfitting is largely suppressed if the ground truth is withheld: the reflection is told whether the prediction was correct and the direction of its error, never the answer itself. Consolidating lessons into skills further discards whatever instance-specific detail remains. Skills learned this way improve performance even from a handful of samples, without the overfitting that few-shot backpropagation exhibits, and gain more as samples grow. 3) Lessons, not only skills, are worth distilling into smaller models. Most prior work neglects both; the few recent exceptions internalize only the skills into a smaller policy, leaving the lessons unused. Yet the instance-specific detail that consolidation discards, of no use at inference time, is a denser training signal than the skill itself: a self-generated trace of how the policy erred and corrected itself. We give a simple baseline that distills both. Each insight is simple, yet all three are routinely overlooked; in our experiments, acting on them consistently improves skill learning and surpasses prior state of the art.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.