acceptodds
Under review as a conference paper at ICLR 2027

Differentiable Harness

Abstract

An agent’s harness defines the instructions, tools, memory, context management, and control flow surrounding its language model. Existing methods improve harnesses by modifying discrete prompt text or agent code, which precludes gradient-based optimization. During generation, however, the model accesses harness text through the internal cache created during prefill. We introduce a differentiable harness: rather than modifying prompt text, we directly optimize the cache it induces, keeping model weights, visible prompt tokens, and tool implementations fixed. We train the cache using teacher-forced replay with cross-entropy and reverse KL regularization against the base distribution, augmented with optional clipped token unlikelihood on error steps. At deployment, an inference engine loads the learned cache artifact at registered token positions without altering the prompt. Across 25 task × harness configurations spanning 16 benchmarks and three frozen model groups (dense, MoE, and hybrid attention/linear-recurrent architectures), differentiable harnesses improve task success rate by up to 20.8% over base agents (with relative gains up to 2.5×) and consistently match or surpass handwritten harnesses. Beyond benchmark accuracy, action-level audits show that the learned cache reduces step-level tool execution errors by 30%, cuts step budget exhaustion by 60% on OSWorld, and avoids the conversational forgetting that degrades weight-adapted baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.