Composition Lives at the Seams: On-Policy Correction for Composing Internalized Skills
Abstract
A skill that lives in a prompt is paid for on every call, and two skill documents pasted into one prompt do not teach a model to chain the skills they describe. Distilling the documents into the weights looks like the cure, and at first it is. A 3B student fine-tuned on flawless demonstrations of 27 documented text skills executes them from memory more reliably than a model over ten times its size reading the documentation. The cure fails where nobody demonstrated anything, on chains one step deeper than the training data. The failure sits at the seams, the handoffs where one skill’s output becomes the next skill’s input; the states that need repairing arise only in the student’s own imperfect attempts, beyond off-policy reach. We therefore let the student attempt compositions, catch the first botched seam with a programmatic verifier, and train on a correct continuation spliced onto the student’s own prefix. This lifts deep chains to the level of a 32B model reading the documentation, at no in-distribution cost. Two controls show what the gain rests on. Freeze the rollouts and it collapses; take corrections from a teacher weaker than the student and it reverses. Composition lives at the seams, in the states where the student stumbles.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.