Passing Is Not Learning: Rethinking Guidance in Coding Agent Training
Abstract
Teacher guidance can turn failed coding attempts into verified successes, but these guided successes do not necessarily yield independent problem solving. We study terminal and repository-repair tasks that a 9B coding agent persistently fails after standard reinforcement learning, asking what guidance leaves behind once removed. Distilling autonomous, hint-assisted, and teacher trajectories improves unaided performance on intervention tasks but yields little gain on held-out tasks when privileged channels are blocked. Audits further show that success filtering also favors trajectories accessing privileged information, while open-channel evaluation can overstate unaided capability. We therefore propose Hint-Twin Coupled Reinforcement Learning (HTC-RL), a two-stage recipe that first distills mixed trajectories under hint-free contexts, then couples reinforcement learning on unhinted rollouts with hint-free online distillation of successful hinted rollouts. The complete recipe achieves 28.1 on Terminal-Bench 2.1 and 57.2 on SWE-bench Verified, 4.5 and 6.4 points above our Stage 0 standard-RL baseline. Against plain RL from the same distilled initialization, its gains extend beyond the initial demonstrations: on 197 terminal tasks absent from those demonstrations but included in continued training, it solves 16 versus 6. On 600 Stage 0-screened hard tasks from repositories excluded from our post-training recipe, it reaches 7.2% unaided 8, versus 5.2% for plain RL. Guidance should therefore be judged by the independent capability acquired from it, not simply by the passes it enables.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.