acceptodds
Under review as a conference paper at ICLR 2027

Your Policy Is Secretly a Harness: The Decision Geometry of Agent Interfaces

Abstract

Language model agents are trained through a harness of prompts, tools, and control logic. Under KL-regularized control, an optimal policy's log odds between executed actions recover the relative action scores that its harness induces, which determine every optimal decision. In web tasks whose control problems we solve exactly, post-training moves a model's action scores toward its training harness's optimum and makes them specific to that harness. Even reversing the order of the observation fields, which keeps every optimal decision, moves these scores and lowers a trained model's success from 85.0% to 19.2%. Action scores fix decisions but leave out the best return a harness offers, which we call its ceiling. An agent's return equals the ceiling minus its mismatch, the return lost to suboptimal decisions along its own trajectories. Training changes only the mismatch. Adding batched clicks raises the ceiling but lowers a trained model's success, and after equal further training under each harness, the batched-click harness gives the best agent. Among swaps between several harnesses, a frozen comparison rejects seven that raise the ceiling, and after adaptation all seven exceed the original agent's return. Ceiling and mismatch measured before adaptation predict the adaptation gains far better than the frozen return change. With the optimum known, mismatch estimated from sampled actions predicts how harness changes and adaptation move return. For coding agents, teacher-based estimates of ceiling minus mismatch order several harnesses by accuracy. A frozen comparison of harnesses therefore mixes what each harness offers with how well the checkpoint uses it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.