acceptodds
Under review as a conference paper at ICLR 2027

Passing Is Not Learning: Rethinking Guidance in Coding Agent Training

Abstract

Teacher guidance can turn failed coding attempts into verified successes, but these guided successes do not necessarily yield independent problem solving. We study terminal and repository-repair tasks that a 9B coding agent persistently fails after standard reinforcement learning, asking what guidance leaves behind once removed. Distilling autonomous, hint-assisted, and teacher trajectories improves unaided performance on intervention tasks but yields little gain on held-out tasks when privileged channels are blocked. Audits further show that success filtering also favors trajectories accessing privileged information, while open-channel evaluation can overstate unaided capability. We therefore propose Hint-Twin Coupled Reinforcement Learning (HTC-RL), a two-stage recipe that first distills mixed trajectories under hint-free contexts, then couples reinforcement learning on unhinted rollouts with hint-free online distillation of successful hinted rollouts. The complete recipe achieves 28.1 on Terminal-Bench 2.1 and 57.2 on SWE-bench Verified, 4.5 and 6.4 points above our Stage 0 standard-RL baseline. Against plain RL from the same distilled initialization, its gains extend beyond the initial demonstrations: on 197 terminal tasks absent from those demonstrations but included in continued training, it solves 16 versus 6. On 600 Stage 0-screened hard tasks from repositories excluded from our post-training recipe, it reaches 7.2% unaided 8, versus 5.2% for plain RL. Guidance should therefore be judged by the independent capability acquired from it, not simply by the passes it enables.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.