EnactOpt: Diagnosis-Guided Harness Evolution for Embodied Agents
Abstract
Embodied agents built on foundation models rely on an execution harness to organize observations, guide actions, and incorporate execution feedback into subsequent decisions. Failures can arise not only from limited model capabilities, but also from how the harness presents evidence and guides its use. These model-external mechanisms provide an adaptation target without updating model parameters, but task-specific interaction experience can be costly to collect. Task-level outcomes alone do not reveal whether relevant evidence was missing from the model input or was available but inadequately used. Moreover, these two aspects are coupled: effective action guidance may require evidence that the current context omits, while exposing that evidence alone does not ensure an appropriate response. We introduce EnactOpt, which uses limited execution feedback to jointly evolve reusable textual skills and context-construction code while keeping the underlying models and runtime fixed. It diagnoses adjacent decisions by relating action feedback to the actual model input and subsequent action. Structured episode summaries guide joint revisions through prompt-level edit operators, and candidate harnesses undergo compatibility checks and validation-based retention. Across VLABench, ALFWorld, and Embodied Agent Interface (EAI), EnactOpt improves task success by approximately 17–41 percentage points over the respective Naive baselines and outperforms the evaluated GEPA, TextGrad, and Trace2Skill baselines. These gains extend to task conditions not used for harness evolution or candidate selection. The results support a complementary route to embodied adaptation: turning limited execution experience into reusable behavioral improvements by jointly refining how evidence is presented and acted upon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.