MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Abstract
Modern agentic systems typically consist of an artificial intelligence (AI) model and a harness, a software layer that manages its control flow and interactions with the environment. Agent performance on long-horizon tasks is often strongly influenced by its harness design. Yet building effective harnesses requires significant human effort due to the combinatorial design space, and this effort must be repeated as models are updated. To automate harness design, existing methods provide limited exploration of this search space. Most optimize only parts of the harness, such as prompts or skills, while others struggle to escape local optima due to their fixed search strategies and exploitative bias in LLM-driven search. In this paper, we present MILO (Meta-evolutionary Island Orchestration), a framework for automated harness discovery that co-evolves the harness and its own search strategy. Three components drive the search: (i) a hierarchical lineage memory of island-based trees that uses rejected mutations as negative evidence to prune unpromising paths and steer toward promising lineages; (ii) per-island mutator agents that combine global search history with feedback on parent weaknesses to rewrite entire harnesses; and (iii) an orchestrator agent that adapts the search by grafting and speciating lineages, reassigning mutators, and revising the curriculum. Together, they make MILO a meta-evolutionary harness-discovery framework that self-adapts its memory, mutators, and curriculum to balance exploration and exploitation. Across three long-horizon benchmarks, Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform eight state-of-the-art harnesses and six search methods with frontier (Opus 4.8) and open-weight (gpt-oss-120b) backbones. With Opus 4.8, it improves resolution rate over its initial harness by %, % and %, respectively, versus the best prior search gains % (GEPA), % (Meta-Harness) and % (no improvement observed). On Terminal-Bench 2.1, it reaches %, above the official leaderboard's top entry (%), while consuming 26% fewer tokens than its initial harness. On EinsteinArena's open problems, MILO-evolved harnesses tighten the best-known upper bounds for Erdős minimum-overlap () and the first and third autocorrelation inequalities (, ).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.