acceptodds
Under review as a conference paper at ICLR 2027

Seek, Ground, Recover: An Agentic Framework for Autonomous Humanoid Loco-Manipulation

Abstract

Autonomous long-horizon humanoid loco-manipulation in open scenes requires task decisions to adapt as locomotion and posture changes alter observations, body-relative target geometry, and action feasibility. Reliable execution requires acquiring sufficient perceptual evidence, converting semantic targets into accurate metric goals, and restoring physical preconditions after failures. We present an agentic framework that couples active evidence acquisition, explicit geometric perception, and execution feedback in a closed decision loop. A vision-language model (VLM) selects observation actions, identifies and verifies task targets, and provides initial target localization. Explicit geometric perception refines these predictions into metric 3D goals for pretrained whole-body policies, separating semantic verification from metric accuracy. During and after execution, the VLM interprets temporal feedback to assess progress and identify failures; the runtime combines these assessments with controller checks to coordinate further observation, recovery, and continuation. Persistent task state retains objectives, perceptual evidence, and execution progress, supporting target re-localization and resumption of interrupted tasks after recovery. A shared goal-and-outcome interface separates semantic coordination from physical control and supports extension of the motor-skill library. Experiments demonstrate the benefits of embodied evidence acquisition and explicit geometric perception for target localization, and characterize recovery under controlled physical failures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.