Attend as You Act: Progress-Guided Adaptive Contextualization for Vision-Language-Action Models
Abstract
Learning from a single video demonstration through in-context learning (ICL) offers a promising way to adapt vision-language-action (VLA) models without task-specific data collection and fine-tuning. We propose *PROACT*, an execution-conditioned one-shot ICL framework that adapts both demonstration representation and behavior grounding. Progress-guided multi-resolution routing dynamically concentrates visual detail and attention on temporally relevant demonstration regions while compressing less relevant context. Predictive behavior grounding further learns a latent operator that links demonstrated behavior to its expected consequences under the current observation, making it actionable in the present scene. Together, these mechanisms transform a fixed video demonstration into a dynamic task context that evolves with execution. Experiments across diverse manipulation settings show that PROACT consistently outperforms approaches that treat the demonstration as static context.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.