acceptodds
Under review as a conference paper at ICLR 2027

Learning In-the-Wild: Selective In-Context WAM Alignment with Internet-Scale Human Video Retreival

Abstract

Human videos offer rich manipulation knowledge, but useful procedures are scattered across long, unstructured recordings. References must show the right operation, and policies must learn what transfers across embodiments. We present WildWAM, which combines RIVER for robot-guided Internet video retrieval with SAIL for selective human-to-robot in-context alignment. RIVER localizes and verifies operations through complementary search routes and event-aware windows; SAIL aligns human context with robot history through Internet pre-adaptation. Across 84 retrieval configurations and seven comparison methods, RIVER achieves 83.3% coverage and 62.0% functional precision. A three-corpus index of 715,565 human intervals supports over 100k annotated human-robot pairs; Open-AoE supplies 66.6% of distinct retrieved human clips. A unified 59,589-episode G1 foundation is trained for 350,000 updates. Internet pre-adaptation reduces downstream action-prediction error by 17.2% and dynamic-region MAE by 4.20% over target-only adaptation. Across five real-robot tasks with 15 trials per task and policy which include both single-arm and bimanual manipulations, WildWAM achieves 44.0% full-task success conditioned on human context, compared with 28.0% for target-only adaptation. These results connect diverse in-the-wild human procedures to reference-conditioned robot learning, demonstrating the potential of learning from the unstructured human videos.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.