HIPPOVLA: Integrating Short-Term Visual Memory into Vision-Language-Action Policies
Abstract
Vision-language-action (VLA) policies conditioned primarily on current observations face temporal ambiguity: similar observations may require different actions depending on prior interactions. We introduce HIPPOVLA, which incorporates short-term visual history through the backbone's existing intermediate visual injection pathways. Its Short-Term Memory Bank (STMB) organizes sparsely sampled historical features into sets corresponding to selected visual levels and camera views, retrieves historical context through cross-attention, and fuses it with current features through learnable gates. The enhanced representations contribute to action prediction without appending historical observation tokens, keeping the language-model input sequence length unchanged. On CALVIN , HIPPOVLA improves average completed sequence length from 4.085 to 4.619 and the success rate for completing five consecutive instructions from 71.2% to 83.0% over the no-STMB baseline with the same backbone and training protocol. Across the four LIBERO suites, it achieves a mean success rate of 98.5%. In real-world repeated pick-and-place manipulation, success rates for completing two and three full cycles both improve over the no-STMB baseline by 30 percentage points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.