PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
Abstract
Long-horizon agent runs generate experience that can improve both the current run and future runs: successful attempts reveal reusable procedures, while failed attempts expose failure modes. Most self-improvement methods process this experience only after a run ends, so they can neither recover that run nor immediately apply and validate the lessons learned, making self-improvement less efficient and less reliable. We argue that self-improvement should instead be live, using emerging experience both to redirect the active run and to update the persistent harness. However, live self-improvement exposes an architectural gap: existing agent architectures do not simultaneously support live correction and a dedicated self-improvement role. Single-agent self-correction can revise the active run, but the same agent must execute the task and judge its trajectory within a limited context, splitting attention between execution and oversight. Subagent delegation separates execution from the main agent, but the main agent typically cannot redirect the active subagent while the subagent is still running. To this end, we present PILOT, a supervisor–worker harness that realizes live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and memory. Across three frozen backbones and four benchmarks, PILOT ranks first in eight of twelve configurations. On Terminal-Bench 2.0, PILOT outperforms counterpart harnesses by as much as 11.8 percentage points. In the self-improvement setting, PILOT gains 14.6 points with GLM-5.1; mean output tokens fall by 42.9%, and successful evaluations per million output tokens rise by 110.3%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.