acceptodds
Under review as a conference paper at ICLR 2027

The Surprising Effectiveness of Few-Step Adaptation for Robotic Manipulation

Abstract

Vision-language-action (VLA) models are pre-trained on large robot datasets and fine-tuned on hundreds of demonstrations per task, yet the natural way to teach a robot something new is to show it once. The prevailing answer for using that one demonstration is in-context learning, which retrains the policy to read demonstrations at inference. We ask whether something far simpler will do: keep the pre-trained VLA as it is and update it with a few gradient steps of its own training loss on the single demonstration. We call this Few-step One-Shot Imitation Learning (FOSIL). It works, and surprisingly well. One gradient step suffices when a familiar skill meets a new object or scene, ten for skills composed in a new order, and thirty for skills the policy could not perform before. We show this on a two-arm robot, with every result backed by 20 physical trials per task over 250 hours of evaluation, and on LIBERO-PRO, where the same demonstration given in context to the same backbone does not improve on the unadapted model. We then ask why so few steps are enough: robot pre-training decides the speed of adaptation, the useful update lives in a single direction per weight matrix, and a demonstration that moves differently from the policy can be refined with the policy’s own motion. A pretrained VLA, it turns out, is already a few gradient steps away from its next task. Physical robot rollouts and videos: https://fosilvla.github.io.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.