acceptodds
Under review as a conference paper at ICLR 2027

BREAD: Transfer-Aware On-Policy Distillation from Boundary Experience

Abstract

Repeated sampling reveals correct programs for tasks that a language model fails to solve greedily. These boundary tasks contrast failed decisions with verified repairs, providing experience for cross-task learning. We introduce BREAD (Boundary Repair Experience for Adaptive Distillation) to connect repair experience with cross-task update utility. BREAD extracts textual repair cards from boundary tasks and uses them to condition a frozen teacher on task-only student rollouts. Temporary updates measure changes in correct-program likelihood across source clusters, and cluster-level utility scores route experience to training receivers. During distillation, candidate AdamW updates start from a shared snapshot and are selected by correctness-probe improvement under retention and displacement constraints. The resulting student generates programs from task-only inputs. Across four model families and four programming benchmarks, BREAD achieves average gains of 7.88 percentage points in macro-average greedy accuracy over the initial models and 1.82 points over the strongest baseline selected separately for each model. Paired analysis separates newly solved tasks from lost solutions, while component and carrier diagnostics examine how repair extraction, utility routing, and update selection contribute to reusable task-only student behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.