ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
Abstract
Recent work extends Recursive Self-Improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve their execution mechanisms from experience. However, achieving and demonstrating generalizable harness RSI remains challenging. First, existing approaches often evolve harnesses directly on evaluation benchmarks or subsets drawn from them, making it difficult to distinguish reusable harness improvements from benchmark-specific adaptation. Second, updates derived from individual trajectories can entangle systematic harness deficiencies with instance-specific reasoning and solution details, leading to task-specific modifications that transfer poorly to unseen tasks. Third, even when recurring behavioral deficiencies are identified, localizing them to the responsible components within a monolithic harness remains difficult. Whole-harness optimization can therefore entangle unrelated mechanisms and produce changes that are difficult to attribute and validate. We propose ModularRSI, a benchmark-disjoint, contrastive, and modular framework for generalizable harness evolution. ModularRSI contrasts successful and failed trajectories for the same task and aggregates evidence across tasks to identify recurring behavioral deficiencies. It further decomposes the evolvable harness into five functional modules: Agent Loop, Tool Use, Observation Management, Context Management, and Task Completion Detection. Each module is evolved independently within a restricted modification scope, after which validated improvements are integrated into a unified harness. To evaluate generalization beyond evolution experience, we additionally curate 2,000 executable evolution tasks from external data sources that are disjoint from downstream evaluation benchmarks. Experiments on TerminalBench 2.0 and SWE-Bench Verified show consistent improvements on unseen in-domain and cross-domain tasks, with the evolved harness also transferring across different foundation models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.