acceptodds
Under review as a conference paper at ICLR 2027

Patches Hide Tests to Write: Deriving Trainable Test-Generation Tasks from Existing SWE Environments

Abstract

Reinforcement learning for repository-level coding agents is limited by the supply of executable tasks, and the work that enlarges it either curates more real issues or manufactures new defects. An issue-resolution instance carries four artifacts: the problem description, the repository snapshot in which the issue is present, the reference patch that fixes it, and the reference tests that the fix turns from failing to passing. Issue resolution withholds the patch and judges with the tests. Exchange the two roles and the same instance defines a second task, writing a test that reproduces the reported behaviour, inside the environment the instance already ships and at no annotation cost. Fail-to-pass status alone is satisfied by tests that exercise no behaviour, such as one that checks whether a line of the buggy source is absent. We therefore require a submitted test to fail both on the buggy snapshot and on snapshots in which the reference patch has been damaged, and to pass both on the repaired snapshot and on versions of it whose source text has been rewritten without changing behaviour. Applied to three collections of instances that differ in where their defects and tests come from, the conversion produces 30,605 fail-to-pass test-generation tasks, of which 23,520 pass an audit of their verifiers. Forty steps of joint reinforcement learning on the two tasks, which draw about 480 tasks in total, raise Qwen3.5-9B from 48.2% to 55.0% on SWE-bench Verified and from 38.5% to 44.5% on TDD-Bench Verified, ahead of training on either task alone. We release SWE-Flip, the conversion pipeline together with the tasks it produced.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.