HERO: Agent Self-Improvement through Harness–Executor Reciprocal Optimization
Abstract
The ability to learn from experience and continually improve is fundamental to intelligence. We introduce HERO, an agent self-improvement framework that adapts experience collection alongside executor learning. A task-conditioned harness generator and an executor, initialized from the same pretrained model, learn through their own environment interactions. Direct preference optimization trains the generator using preferences derived from verified execution outcomes. Successful action trajectories provide supervision for executor fine-tuning, with student inputs reconstructed under a fixed deployment harness. The updated executor then supplies feedback for subsequent generator adaptation. This alternating process uses self-generated experience without additional teacher-provided solutions or critiques. We evaluate HERO with Qwen3.5-9B on ALFWorld and WebShop. On ALFWorld, the final executor reaches 97.0% success on the official unseen split. On WebShop, three consecutive DPO–SFT stages improve success on the official 500-task test split from 31.2% to 45.6%. The final executors achieve these results under fixed deployment harnesses, requiring no extra generator calls. Analysis shows that the trained generator yields more clean successful trajectories and shorter average rollouts than its untrained counterpart. These findings highlight HERO as a way to shape the experience from which an executor learns.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.