Seek, Ground, Recover: An Agentic Framework for Autonomous Humanoid Loco-Manipulation
Abstract
Autonomous long-horizon humanoid loco-manipulation in open scenes requires task decisions to adapt as locomotion and posture changes alter observations, body-relative target geometry, and action feasibility. Reliable execution requires acquiring sufficient perceptual evidence, converting semantic targets into accurate metric goals, and restoring physical preconditions after failures. We present an agentic framework that couples active evidence acquisition, explicit geometric perception, and execution feedback in a closed decision loop. A vision-language model (VLM) selects observation actions, identifies and verifies task targets, and provides initial target localization. Explicit geometric perception refines these predictions into metric 3D goals for pretrained whole-body policies, separating semantic verification from metric accuracy. During and after execution, the VLM interprets temporal feedback to assess progress and identify failures; the runtime combines these assessments with controller checks to coordinate further observation, recovery, and continuation. Persistent task state retains objectives, perceptual evidence, and execution progress, supporting target re-localization and resumption of interrupted tasks after recovery. A shared goal-and-outcome interface separates semantic coordination from physical control and supports extension of the motor-skill library. Experiments demonstrate the benefits of embodied evidence acquisition and explicit geometric perception for target localization, and characterize recovery under controlled physical failures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.