HARVEST: Harness-Guided Trajectory Learning for Self-Evolving Language Models
Abstract
Verified self-training improves language agents by training on their own trajectories with correct final answers, but its progress is limited by the trajectories the current agent can already generate. We observe two structures in the remaining failures. First, many failed trajectories break at a local step that the same agent has handled successfully in other questions. Second, question difficulty depends on the agent, so improvement can unlock new correct trajectories from previously unsolved questions. We introduce HARVEST, a self-evolving training framework built around these observations. A rescue harness targets current failures by retrieving successful transitions from the agent's own verified trajectories and falls back to answer-conditioned generation when retrieval fails. Only trajectories with correct final answers are retained, and all inserted guidance is removed before training. The first iteration combines rescues with a diverse anchor of the initial agent's verified trajectories, while later iterations combine fresh rescues with newly unlocked correct trajectories from the updated agent. Across three backbones, HARVEST improves over the strongest baselines by 7.4, 4.2, and 3.9 percentage points on ToolEvo, ToolQA, and MATH. At equal generation budget on MATH, ordinary resampling misses 64% of the questions rescued by HARVEST. Across later iterations, newly unlocked trajectories increasingly replenish the training data, while combining them with fresh rescues gives the strongest improvement. These results show how recovering local failures and learning from progressively unlocked trajectories can sustain agent self-improvement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.