Harness Evolution Hits a Ceiling: When Weight Training Should Begin
Abstract
Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We follow one line through both: let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. We show that the right lever can be read off the agent's failure composition. Labelling failed trajectories by the first signal that fires separates process failures (blocked calls, loops, exhausted step budgets) from content failures (a delivered plan that is poor). Harness evolution repairs the former, the behaviour it instils can be trained into the weights, and content failures are what weight training is for. On DeepPlanning, a verifiable planning benchmark with a held-out split, a self-evolving harness loop, in which an LLM proposes edits to its own harness and a judge verifies them, lifts the held-out score of Qwen3.5-4B from 0.16 to 0.30 and of Qwen3.5-9B from 0.32 to 0.44, and the gain lands where the claim places it: for 4B, held-out delivery rises from 55% to 90% and loops fall, while content failures are left for the weights. Ablations pin the gain on capability-granting components rather than on enforcement, and the loop's own evidence package shows why it stops where it does. LoRA adapters trained on evolved-harness trajectories then internalise the gain: under the original harness they add +0.13 on held-out tasks for both sizes; on 4B they stack with the harness to more than double the held-out score, and on 9B the adapter alone matches the full evolution line, cutting content failures from a quarter of trajectories to one in twenty. A placebo adapter trained on answer-shuffled trajectories falls below the base model, so the gain is in the trajectory content; full fine-tuning on a larger trajectory set adds a further +0.04. The same loop transfers to WebArena-Lite (+0.09 on 117 unseen tasks), where its gain lives in what the model sees, two pages of history instead of one, and adapters trained on those trajectories do not add to it. The result is a diagnose-then-intervene rule applied twice: read the failure composition, and evolve the harness for process failures and train the weights for content failures; then read what the accepted edits changed, and train in the gains that changed what the model writes while keeping the harness for the gains that changed what it sees. Scores are four-rollout means against fresh anchors and comparisons are drawn within one night, except where marked, across eight models from six families and two benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.