acceptodds
Under review as a conference paper at ICLR 2027

What Makes Good In-Context Examples for Vision-Language-Action Models?

Abstract

Vision-Language-Action models (VLAs) have demonstrated strong capabilities across a wide range of physical manipulation tasks. To enable more versatile generalization, recent studies have sought to improve their In-Context Learning (ICL) ability, i.e., to perform new physical tasks by conditioning on a few related demonstrations at inference time. However, what information in these demonstrations actually drives ICL performance remains poorly understood. Prior work largely focuses on demonstrating ICL gains, leaving the roles of different context components underexplored. In this paper, we systematically investigate what makes in-context demonstrations effective for VLAs. We find that the action tokens in the demonstrations play the most important role, whereas removing images and language results in only a marginal decrease in performance. Our analysis also reveals a surprising fact: the pairing between observations and actions is largely unnecessary. Given a selected context, the correct mapping between images and action tokens is actually not required to perform ICL. Shuffling the ground-truth action tokens among demonstrations in the context window has little effect on the performance. We further find that ICL becomes more beneficial as tasks require greater generalization, while increasing the number of demonstrations can eventually hurt performance. Together, these findings provide practical insights into effective context construction and raise fundamental questions about how current VLAs use demonstrations for in-context learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.