TEST-TO-TEACH: EXECUTION-CONDITIONED ONPOLICY DISTILLATION FOR AUTONOMOUS CODINGAGENTS
Abstract
Unit tests are widely used to evaluate and refine code generated by autonomousagents, but their outcomes are often reduced to binary rewards or repair prompts.We introduce Test-to-Teach, an execution-conditioned, on-policy distillationframework that uses program behavior as supervision. The student generates code,tests, and repair trajectories. The framework then constructs counterfactual testsby perturbing variables, inverting boundary conditions, substituting dependencies,and contaminating state. These tests target implementations that pass the existingsuite for the wrong reason. Without seeing reference solutions, the teacher derives local behavioral constraints, causal failure graphs, and minimal repair principles from execution traces. Execution information gain determines the weightof each distillation example. Under a matched training budget on SWE-benchVerified, Test-to-Teach improves the resolved-task rate from 34.8% to 43.6% overstandard trajectory distillation trained on the same teacher traces. Across threecross-repository benchmarks, it raises the hidden-test pass rate by 9.7 percentagepoints and reduces patch overfitting by 31.2%. In online repair on HumanEval+,it reaches a solved-task rate of 86.59%. The results show how execution evidencecan support code generation, failure attribution, and repair planning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.