acceptodds
Under review as a conference paper at ICLR 2027

Beyond Rejection Sampling: Learning from Guided Rollouts to Improve LLMs for Software Engineering

Abstract

Rejection-sampling supervised fine-tuning (SFT) for coding agents retains only trajectories that pass executable tests. On difficult repository tasks, repeated unguided rollouts often produce no successful trajectory. Consequently, **many available tasks and execution environments yield no retained SFT examples**, wasting executable software-engineering data. We use answer-derived guidance to recover training trajectories. Based on when guidance enters an interactive rollout, we design three mechanisms: **Pre guidance** derives a hint from the reference solution before initial rollout; **Post guidance** translates an unsuccessful attempt and the reference solution into a hint for fresh retry; and **Intra guidance** provides answer-derived assistance at eligible interaction steps. Our method, *SFT with Guided-Success Augmentation* (GS-SFT), augments rejection-sampling SFT with verifier-accepted guided trajectories and uses ordinary supervised fine-tuning. We distinguish **Supply**, retained SFT examples normalized by the number of executable tasks, from **Utility**, the task resolved rate after SFT, measured on SWE-bench Verified and SWE-bench Multilingual. Across the ordered guidance constructions, Supply increases steadily, whereas **Utility peaks at an earlier or intermediate non-oracle condition and then decreases under more explicit guidance**. Direct reference patches produce the most data but not the highest resolved rate. For offline guidance, our results favor **retaining hints in the SFT input rather than removing them**. The additional guided data produces the largest resolved-rate gains on issues with no actionable solution or reproduction lead, while also improving issues that already identify a likely solution or code location. Overall, moderately informative answer-derived guidance provides a simple way to use executable software-engineering data more efficiently during SFT.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.