BoxBox: Recursive Harness Self-Improvement for Long-Running Performance Optimization
Abstract
Automated harness optimization improves agents by learning from trajectories and task feedback. Existing offline methods use reference answers to optimize harnesses for reuse on held-out tasks. However, many real-world tasks lack reference solutions, while long runs expose new difficulties that offline-optimized harnesses may not address. To study harness optimization in these settings, we introduce , a benchmark for long-running performance optimization of software projects without reference solutions, and further investigate whether agents can use task-specific experience to recursively improve their own harnesses during an ongoing task. To this end, we introduce , which enables agent-triggered harness revision and kernel-managed hot reloading while preserving execution context. We evaluate on under matched test-time budgets, comparing it with Frozen Harness, offline Harness Evolution, and Harness Scaling. achieves higher speedup with more efficient use of task time and model cost, reaching a 4h of \(1.70\times\), compared with \(1.54\times\) for the strongest baseline. Code is available at https://anonymous.4open.science/r/BoxBox-1C35/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.