acceptodds
Under review as a conference paper at ICLR 2027

Beyond Context Sensitivity: Cross-Embodiment In-Context Policy Learning

Abstract

Adapting robot policies to new tasks through post-training incurs demonstrationcollection and optimization costs. In-context learning (ICL) avoids test-time weight updates, but robot-reference methods still require demonstrations from the same embodiment. We ask whether human videos can provide references without robot teleoperation. Human-motion retargeting often relies on hand-designed kinematic correspondences, tying motion transfer to a robot-specific mapping. We introduce MIRAL (Mixed-embodiment In-context Reference Action Learning), a cross-embodiment ICL pretraining framework with a shared human–robot action interface. Disjoint action coordinates and modality-specific reference-action projectors support joint learning with human and robot references and queries. To discourage reference neglect, Context Error Margin (CEM) training fits query actions under matched references and raises mismatched prediction error. On The Imitator Game, we achieve 35% zero-shot success, 22 percentage points above the strongest listed comparator. Real-robot experiments on soybean pouring and tableware arrangement demonstrate effective human-reference-conditioned control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.