Meta²: Precomputing Reusable Knowledge for Online Harness Evolution in Long-Horizon Agents
Abstract
Harness optimization improves long-horizon agents by refining the prompts, tools, memory, and hooks around a frozen model. Offline methods can learn from completed trajectories, but deploy one shared harness that cannot react to evidence revealed during a new run. Online evolution can react within the run, but may have to reason about every intervention from scratch. We introduce Meta², which bridges these regimes by precomputing reusable knowledge about how to evolve a harness, rather than precomputing one universal harness. Offline, Meta² compares three rollout conditions: the task agent alone, the current Meta-Skill-guided system, and a training-only evolver with privileged verifier information. These contrasts provide evidence that an evolver should act, hold, or avoid an intervention, which is aggregated across tasks into a textual Meta-Skill M. Online, M is frozen. At checkpoints within an uninterrupted held-out run, the evolver combines it with active-run evidence and chooses Keep or a task-local Evolve edit; the task agent then continues from the same environment state. Neither held-out verifier feedback nor task-specific edits propagate across runs. Meta² achieves the highest point estimates among the evaluated methods on both Terminal-Bench 2.1 and DeepSWE 1.1 with GPT-5.6-Terra, reaching 81.00% and 44.44%, respectively. These scores are 6.33 and 5.33 percentage points above the strongest comparators. With DeepSeek-V4-Flash-0731, Meta² achieves 77.67% on Terminal-Bench 2.1, 3.00 percentage points above the strongest comparator. Cumulative ablations associate the largest additional gains with contrastive Meta-Skill construction, while checkpoint sensitivity exposes a responsiveness–cost tradeoff. These results support precomputing reusable evolution knowledge while keeping each realized harness edit local to the active task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.