AgentSpec: Evolving Embodied-Agent Harnesses through Typed Composition
Abstract
Embodied LLM agents run inside harnesses of perception, memory, reasoning, reflection, and action, and coding agents can now write and revise these harnesses. We present AgentSpec, a typed specification and runtime in which each component has a declared contract and a complete harness is an executable, diffable program that a coding agent can modify, compose, run, and trace. With it we compare how one coding agent constructs harnesses for one fixed executor: evolving reasoning, memory, and reflection components inside declared contracts over five rounds of execution feedback, writing a free-form program once, and selecting a library configuration, on DeliveryBench, ALFRED, MiniGrid, RoboTHOR, and DeliveryGym. On DeliveryBench the evolved harness, which compresses the action history into an outcome-aligned ledger and adjudicates between sampled decisions, exceeds both baselines, and its advantage persists on average when the frozen harness is re-executed. On MiniGrid the one-shot programs, which decode the rendered observation into a map and plan on it, exceed the evolved harness by a wide margin on the development suite and on held-out tasks. Reading the generated implementations against their execution records shows that, within the editing scope of this study, the two workflows divided work differently between model and program: evolution refined what the executor saw at each step and kept it in the loop, whereas free-form generation moved state maintenance into code and consulted the executor far less often, at the same recorded success rate on ALFRED and a fraction of the cost. AgentSpec makes such divisions of computation inspectable, attributable to specific component changes, and accountable in design, evaluation, and deployment cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.