acceptodds
Under review as a conference paper at ICLR 2027

SWE-OPD: Enhancing Agentic Coding via On-Policy Distillation

Abstract

In this paper, we present our practical experience with on-policy distillation for enhancing agentic coding capabilities. Specifically, we leverage a large and capable teacher model and successfully transfer its agentic coding capabilities to a smaller student model through vanilla on-policy distillation. Unlike prior work that primarily focuses on mathematical reasoning and relatively simple agentic tasks such as web search, we extend the application of on-policy distillation to the more challenging domain of agentic coding. We conduct extensive experiments on various datasets and consistently demonstrate its effectiveness on various SWE benchmark. For example, on-policy distillation improves the performance of Qwen3.5-9B on SWE-bench Verified from 57.9% to 62.8%, around 5% enhancement under the supervision of Qwen3.6-27B. Additionally, to alleviate the substantial GPU memory overhead, we develop an offline variant of on-policy distillation that substantially reduces memory consumption with only a slight performance degradation. Finally, we show some intriguing findings during our training and evaluation and make our code available, hoping to further push the development of on-policy distillation on agentic coding with community.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.