Weights or Words? Where to Spend an Agent's Adaptation Budget
Abstract
An agent can be improved by training its weights, or by rewriting the text wrapped around it: the system prompt, the tool descriptions, and the accumulated notes that tell it how to behave, which we call its harness. The two literatures disagree about which of these is the better use of a fixed budget. Reflective prompt rewriting has been reported to beat GRPO using an order of magnitude fewer episodes; on longer-horizon agent tasks, however, the same family of methods has been reported to lose to a single hand-written prompt. Yet neither comparison holds compute fixed, and only under a fixed budget does the comparison yield a fair and usable conclusion. We therefore hold it fixed, comparing training the weights alone, rewriting the harness alone, and doing both, under identical FLOPs on a multi-turn retrieval agent. Our systematic experiments reveal that rewriting the harness does not turn budget into accuracy: a tenfold budget increase changes success by at most , and every interval contains zero. The same increase spent on the weights, by contrast, yields . This asymmetry is not specific to one setting: the null for rewriting reappears in a second environment and at a second model scale, whereas the advantage of doing both, the strongest configuration on retrieval, does not survive the change of environment. Nor do the two sides combine, since transplanting a rewritten harness onto separately trained weights adds nothing; they are not interchangeable stores of the same knowledge. The null itself resists the obvious explanations, as a rewriter nine times larger writes no better prompts, while a third of the proposals are accepted and read like sound advice. Taken together, these results suggest that what limits such an agent is not the quality of the harness it is given but its capacity to exploit it, which rewriting cannot supply.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.