acceptodds
Under review as a conference paper at ICLR 2027

RoboICL: Embodied In-Context Learning with GPT-6 Astra

Abstract

In-context learning enables LLMs to adapt to new tasks from demonstrations without updating model parameters. We study whether the same principle can support robot control, where demonstrations must connect language and vision to continuous actions and subsequent decisions must account for execution outcomes. provides a general-purpose multimodal starting point, but its zero-shot performance varies substantially across manipulation tasks. We introduce RoboICL, a framework for closed-loop control that combines a static prefix of compressed expert demonstrations with a bounded dynamic context of the robot's own interactions. Both contexts use the same observation–action call–execution receipt–next observation format, allowing the model to distinguish proposed actions from those actually executed. Bounded memory retains temporally distributed interactions and recent feedback while marking omitted intervals. At each decision round, predicts Cartesian end-effector action blocks from the current observation and retained context, without task-specific model updates or an auxiliary learned action policy. In RoboDojo simulations, adding demonstrations is associated with gains on the evaluated shot-study tasks. Across three single-arm real-robot tasks, mean normalized progress rises from 14.45 at zero shot to 78.89 with three demonstrations. A separate exploratory study coupling JEV with RoboICL substantially improves inference speed on RoboDojo while preserving control capability. These results show that demonstrations and execution-grounded interaction history can serve as an effective inference-time interface for direct robot control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.