Learning to Self-Train: Recursive Self-Improvement for Coding Agents
Abstract
Learning to self-train requires a coding agent to diagnose failures, construct useful practice, and solve the resulting tasks. We introduce AI4AI, a teacher-assisted framework that trains these behaviors together through recursive supervised fine-tuning. A single student model serves in three prompt-conditioned roles: a Critic diagnoses capability gaps from execution traces, a Producer constructs repository-grounded practice tasks, and a Solver generates candidate repairs. Execution-based admission checks filter the generated supervision, and joint supervised fine-tuning updates all three roles. Successive updates therefore change both the agent and the supervision it helps generate. Teacher demonstrations initialize the process, followed by increasing student contributions within each role. Across eight training requests starting from a 4B model, the 500-task resolution rate rises from 7.2% at the original base model to **20.2% after cold start and 62.2% at the final request**, a gain of **42.0 percentage points** beyond cold start. In controlled role evaluations, Critic diagnosis accuracy rises from **39.0% to 78.5%** on 200 held-out failure traces, and Producer valid-task yield rises from **7.0% to 31.5%** over 200 proposals per model. With the R8 Solver fixed, R8 Critic feedback raises retry success from **19.0% without feedback to 33.5%** on 200 failed tasks. The results show improvements in coding, diagnosis, and task generation, with Critic feedback providing a measurable benefit to repair.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.