InterGuide: Geometric Guidance for Video-based Loco-Manipulation Learning
Abstract
Scaling imitable data is critical for learning general-purpose humanoid loco-manipulation, and in-the-wild human–object interaction (HOI) videos offer an economical and virtually unbounded source. However, motion reconstructed from such videos suffers from depth ambiguity, occlusion, temporal jitter, and often lacks detailed hand motion, making the resulting reconstructions difficult to imitate reliably via reinforcement learning (RL). We present InterGuide, a framework that leverages geometry-derived interaction guidance to refine HOI reconstruction, and formulate hand-specific rewards to facilitate imitation learning of fine-grained loco-manipulation activities. This guidance takes two complementary forms: pre-sampled coarse gripper poses produced by DenseGripper for grasp interactions, and sparse surface contact anchors for non-grasp interactions (e.g., pushing, supporting). It is integrated across three stages. First, a template-free reconstruction pipeline associates these geometric hypotheses with multi-frame video evidence to refine hand–object contact, depth ordering, and relative scale. Second, interaction-preserving retargeting maps heterogeneous human shapes onto a unified humanoid, refining the palm with the selected gripper pose while aligning gravity and ground support. Third, phase-conditioned grasp rewards use the coarse gripper poses to guide approach, acquisition, and release, while relaxing per-finger tracking to permit contact-adaptive finger motion; non-grasp interactions instead follow the refined contact anchors. Across diverse sequences spanning varied object categories, InterGuide reduces reconstructed contact error and delivers substantial gains in hand-motion imitation success rate and grasp fidelity over competitive baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.