acceptodds
Under review as a conference paper at ICLR 2027

BoxBox: Recursive Harness Self-Improvement for Long-Running Performance Optimization

Abstract

Automated harness optimization improves agents by learning from trajectories and task feedback. Existing offline methods use reference answers to optimize harnesses for reuse on held-out tasks. However, many real-world tasks lack reference solutions, while long runs expose new difficulties that offline-optimized harnesses may not address. To study harness optimization in these settings, we introduce , a benchmark for long-running performance optimization of software projects without reference solutions, and further investigate whether agents can use task-specific experience to recursively improve their own harnesses during an ongoing task. To this end, we introduce , which enables agent-triggered harness revision and kernel-managed hot reloading while preserving execution context. We evaluate on under matched test-time budgets, comparing it with Frozen Harness, offline Harness Evolution, and Harness Scaling. achieves higher speedup with more efficient use of task time and model cost, reaching a 4h of \(1.70\times\), compared with \(1.54\times\) for the strongest baseline. Code is available at https://anonymous.4open.science/r/BoxBox-1C35/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.