acceptodds
Under review as a conference paper at ICLR 2027

EmbodiedHarness: Adaptive Harness Optimization for Self-Evolving Embodied Agents

Abstract

Embodied agents operating in partially observable, long-horizon environments depend on a model-external harness that organizes skills, manages context, verifies completion, and handles failures. Improving this harness from sparse feedback requires deciding not only how to revise a component, but which component to revise. We introduce EmbodiedHarness, a non-parametric framework that adapts these four components from environment interactions while keeping planner and executor weights frozen. The same frozen vision-language model acts as both planner and optimizer in a diagnose-revise-validate loop. To guide component selection, we propose anchored counterfactual replay: restoring simulator state and recorded history at an observable deviation anchor, applying a bounded single-component revision, and comparing candidate and parent outcomes through paired suffix rollouts. Selected revisions undergo paired validation and are accepted only when they improve observed success without losing any parent-solved validation episodes, subject to cost and catastrophic-event constraints. Across three simulated embodied benchmarks, harness optimization improves held-out performance. On VLABench, EmbodiedHarness raises average success from 34.8% to 48.9%; in a separate comparison using SHAPER's backbone, it exceeds SHAPER by 10.2 percentage points over the four jointly reported capability categories. On ESI-Bench, micro accuracy increases from 39.8% to 62.8%. On EmbodiedBench, it achieves 62.0% on Habitat and 65.3% on Navigation, exceeding the reported train-free baselines using the same backbone. Ablations support the contributions of component selection, anchored replay, and validation to effective harness adaptation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.