acceptodds
Under review as a conference paper at ICLR 2027

Improving Test-Time Scaling for Software Engineering Agents

Abstract

Test-time scaling (TTS) offers a promising path for solving complex software engineering tasks, yet its effectiveness depends on generating correct solutions within a fixed rollout budget. We introduce EntroPO, an entropy-enhanced preference optimization framework for multi-turn, tool-using agents. We formulate agent alignment as an entropy-regularized MDP and derive the EntroPO-DPO and EntroPO-KTO objectives. Theoretically, we analyze why standard DPO/M-DPO objectives collapse and show that EntroPO promotes both accuracy and diversity among correct solutions. Empirically, we validate EntroPO on SWEBench, LiveCodeBench, and TerminalBench, where it achieves state-of-the-art results among open-weight models. At the time of leaderboard submission, our 30B model ranked 1st on SWEBench-Lite and 4th on SWEBench-Verified among open-weight submissions, trailing only models over 10× larger on SWEBench-Verified. These results demonstrate the effectiveness of EntroPO for improving software engineering performance with test-time compute scaling. To support further research, we release our code and data through an anonymous link.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.