Video or Action? Demonstration Modality Matters in Embodied In-Context Learning
Abstract
Demonstrations can guide an agent’s predict proper actions through In-Context Learning (ICL) without updating model parameters. Yet identical demonstration can be represented in different modalities (video or action) while providing different task information (e.g., in videos, movements can accompany environmental responses). In this work, we study the effectiveness and complementarity of embodied demonstration modalities. Theoretically, we assume that demonstrator and learner responses depend on a shared, unobserved task condition through distinct feature maps. Then, we derive necessary and sufficient conditions for exact learner prediction in the infinite-context limit. These conditions distinguish insufficient task information in the demostration from insufficient representation capacity: increasing representation width cannot recover task-relevant information absent from the demonstration. Practically, we evaluate policy execution in various 2D and 3D environments as well as future video and action prediction on aligned human–robot recordings. Independently trained modality variants measure the benefit of each input configuration, while fixed-weight interventions test reliance on demonstration content. Our experiments yield three key findings. (i) Video alone supports effective execution, and action trajectories can achieve comparable performance when the demonstrator adapts its actions to environmental conditions. (ii) Video and measured action trajectories provide complementary cues for human–robot prediction: together they yield lower state and visual errors than either alone. Gains from combining modalities in policy execution vary across tasks. (iii) Incorrect demonstrations are more harmful than absent ones, but correct demonstrations can rescue drifting predictions even when supplied after prediction begins, without restarting the rollout or updating model weights. Code is available at anonymous.4open.science/r/anonymous-core-code-CBD2.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.