acceptodds
Under review as a conference paper at ICLR 2027

TEST-TO-TEACH: EXECUTION-CONDITIONED ONPOLICY DISTILLATION FOR AUTONOMOUS CODINGAGENTS

Abstract

Unit tests are widely used to evaluate and refine code generated by autonomousagents, but their outcomes are often reduced to binary rewards or repair prompts.We introduce Test-to-Teach, an execution-conditioned, on-policy distillationframework that uses program behavior as supervision. The student generates code,tests, and repair trajectories. The framework then constructs counterfactual testsby perturbing variables, inverting boundary conditions, substituting dependencies,and contaminating state. These tests target implementations that pass the existingsuite for the wrong reason. Without seeing reference solutions, the teacher derives local behavioral constraints, causal failure graphs, and minimal repair principles from execution traces. Execution information gain determines the weightof each distillation example. Under a matched training budget on SWE-benchVerified, Test-to-Teach improves the resolved-task rate from 34.8% to 43.6% overstandard trajectory distillation trained on the same teacher traces. Across threecross-repository benchmarks, it raises the hidden-test pass rate by 9.7 percentagepoints and reduces patch overfitting by 31.2%. In online repair on HumanEval+,it reaches a solved-task rate of 86.59%. The results show how execution evidencecan support code generation, failure attribution, and repair planning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.