HumanoidHarness: Self-Improving Humanoid Task Execution through Experience-Guided Agentic Policy Revision
Abstract
Pretrained robot policies provide reusable action capabilities, but completing humanoid tasks requires maintaining situational awareness, adjusting body–target alignment, and adapting execution to the environment. Whole-body motion changes these conditions throughout an interaction, making feedback and recovery essential to coordinating learned capabilities. We introduce HumanoidHarness, a self-improving agentic framework that organizes pretrained controllers through an editable execution policy. The policy specifies controller stages, observation checks, and recovery transitions, and is validated and compiled into a state machine. During execution, the harness bounds each action segment and uses new observations to assess progress and select the next state. Across attempts, a reflection journal and task memory retain execution evidence to guide policy revision and reuse. This design separates immediate execution feedback from experience-guided changes to the policy, while keeping pretrained controller parameters fixed. The same agent logic supports simulation and hardware using egocentric vision, proprioception, and task instructions. Results on HumanoidArena show improved task completion and a generally steady increase in success rate across policy revisions with episode memory, while hardware examples illustrate how coordinated actions can adapt to environmental constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.