TAME: Transfer-Aware Model-Harness Co-Evolution through Learned Harness Revisions
Abstract
Agentic post-training relies on an execution harness that both assists the model and shapes the trajectories from which it learns. A harness that improves immediate success, however, need not produce capability that transfers when that support changes. We examine this distinction through crossed training- and deployment-harness evaluation, separating same-harness performance, capability retained after support removal, and transfer across a common evaluation family. This evaluation tests for harness rank reversal, in which a harness with stronger initial performance yields a weaker trained model under common deployment conditions. This motivates TAME (Transfer-Aware Model–Harness Co-Evolution), which trains a Feedback-Agent language model in the self-evolution framework with reinforcement learning to generate persistent, executable revisions to the training harness. Each revision precedes a fixed-budget inner training stage for the Task-Specific Agent and receives a reward based on the resulting performance change under reference harnesses. We employ token-level proximal policy optimization to train the Feedback-Agent on complete revision-and-training trajectories, assigning credit through transfer gains across subsequent model updates. This makes teaching value an explicit objective for harness adaptation: the capability the model acquires through training, rather than only the success it achieves while supported by the harness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.