acceptodds
Under review as a conference paper at ICLR 2027

Harness Learning Enables Generalizable Test-Time Adaptation

Abstract

A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Different tasks call for different ways of organizing these operations, creating a need to adapt the harness using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver’s harness using execution information. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer through reinforcement learning using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, while both models remain frozen. Experiments across reasoning and multi-hop question answering show improved revision quality and transfer to unseen tasks. Policies trained on individual revisions can continue improving harnesses over multiple rounds, while the benefits of training on revision sequences vary across settings. These findings suggest a path towards continually learning agents that turn accumulated experience into generalizable improvements.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.