acceptodds
Under review as a conference paper at ICLR 2027

GICAT: Goal-conditioned In-Context Action Translation for Generalizable Robot Manipulation

Abstract

General-purpose robots need to interpret diverse task goals, anticipate feasible changes in the environment, and realize those changes through embodiment-specific control. Video generation models provide a promising action-free planning interface, but their planning predictions cannot be directly executable by robots. Therefore, we present GICAT, a decoupled closed-loop framework that separates video planning from action generation. Given a current image and task instruction, a video generation model in GICAT predicts a short-horizon visual future. Unlike frame-by-frame action translation, an in-context action translation part of GICAT conditions a flow-matching model on the current image, goal image, robot state, and retrieved in-context demonstrations to generate executable action chunks. By training the action translation model as an in-context learner, we can apply it to unseen tasks during the inference stage without any weight updates. During execution, GICAT follows the predicted action chunk in a receding-horizon loop and replans from real environment observations. GICAT outperforms decoupled video-to-action baselines on Meta-World and ManiSkill and shows stronger cross-task generalization to unseen tasks during in-context action translation training. These results highlight the potential of decoupled video planning and in-context action learning as an effective route toward generalizable video-model-based robot control.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.