SNaP: Self-Narrating Policy with Fast-Weight Context for VLA Control
Abstract
Long-horizon robot manipulation requires recalling past observations and tracking task progress that may not be apparent from the current observation. We introduce SNaP, a self-narrating vision-language-action (VLA) policy that builds context from its own past narrations to guide future decisions. The policy generates short descriptions of its current subtask and relevant task state, using them to condition the current action and accumulating them as context for future narrations and actions. Rather than retaining a growing text history, SNaP encodes past narrations into token-level associations stored in a fixed-size fast-weight narration bank that is updated online during deployment. A separate sensory bank retains proprioceptive context. Readouts from both banks condition subsequent narration and action prediction. The policy thus builds and uses its context online from its own descriptions and observed states. Across three real-robot manipulation tasks requiring object-location recall, cue-based counting, and multi-stage progress tracking, SNaP achieves 90.0% average success, outperforming the strongest evaluated baseline by 23.3 percentage points. Targeted context interventions further show that omitting specific past events changes subsequent subtask predictions, supporting the policy's use of accumulated context.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.