acceptodds
Under review as a conference paper at ICLR 2027

TMT: Co-Evolving Models and Harnesses with Local Test Feedback

Abstract

The performance of a coding agent depends on both the model and the harness. A complete task rollout may succeed despite ineffective actions, or fail despite correctly implemented subtasks. We propose Task–Method–Test (TMT), a framework that uses test feedback to improve both components. TMT splits each original task into subtasks and, following test-driven development (TDD), generates corresponding tests for them. The tests are admitted and fixed before implementation. Each subtask is then addressed by one or more methods. Each method comprises one or more model responses and their associated tool calls. During rollouts on training tasks, TMT records these methods and their test outcomes. Records whose tests pass are used for model fine-tuning, including records from rollouts that fail whole-task evaluation. Records whose tests fail guide harness adaptation through aggregation, diagnosis, and targeted repair, even when their rollouts pass whole-task evaluation. We evaluate the first TMT update cycle, in which the model and harness are each updated once, using a Model × Harness matrix on SWE-bench to measure the individual and combined effects of these updates. We then assess the updated system on eight zero-shot coding benchmarks. The updated model–harness pair outperforms the vanilla system and the systems with only a model update or only a harness update in whole-task pass rate, while using fewer rollout tokens. The updated harness generalizes well to frozen external models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.