Latent Harness: Compiling Agent Behaviors into Post-Block Activation Shifts
Abstract
Harness engineering and self-evolving agents improve language-model agents quickly, yet their gains live outside the model: the harness is re-sent and re-paid on every call, and the standard route to internalization is a weight update that is costly to iterate and fixed once merged. We propose compiling the harness instead. Latent Harness distills an external condition, either an explicit process-level behavior instruction or the behavioral contract of an agentic harness, into a rank-8 activation shift applied at the output of every decoder block of a frozen backbone (0.03–0.05% of parameters). At inference the compiled bundle carries no harness text, adds zero tokens, and exposes three runtime controls: injection strength, injection phase, and shift-space mixing. Compiled behaviors recover the source-harness gain in 8 of 10 preregistered cells, and where the explicit source pays the heaviest per-call token tax the compiled form surpasses it outright; they transfer zero-shot across tasks (5 of 7 pairs positive), compose in shift space while stacked text instructions interfere destructively, and lift budget-matched harness search by 1.45. On three agentic benchmarks the same operator aligns policy to the harness contract: under matched data, capacity, and effective strength it outperforms same-capacity LoRA, which needs 11.1 the parameters to reach the same band, and it regresses nothing where the base policy is already strong. Across 1B-to-8B models the method skeleton is invariant while gain locations move with the model-task pair, which delimits where compiled harnesses should be deployed. Our code is available at https://anonymous.4open.science/r/latentharness-C574.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.