acceptodds
Under review as a conference paper at ICLR 2027

Deconfounding and Compressing History Memory in Streaming Vision-Language-Action Policies

Abstract

How much explicit history does a streaming vision-language-action (VLA) policy need? A common test shortens its attention window, but this intervention can also shrink the physical key/value (KV) pool and change which history remains visible during denoising. We separate these quantities in LingBot-VA and show that an apparent collapse from 97.5% to 0% success on LIBERO-Long is a cache-capacity confound: prediction slots displace the previous chunk’s video tokens. With the window and pool controlled, retaining the immediately preceding chunk achieves 99% success versus 97% with full history; a three-task RoboTwin replication shows the same zero-to-one retention transition. Content interventions test what the retained chunk contributes. We then combine short retention with KV quantization. On LIBERO-Long, two chunks at K4/V2 achieve 99% success, with a paired difference of 0 percentage points (95% CI: [−3, 3]) from the native W = 30 FP reference. Their history-KV storage equivalent under simulated quantization is 28.5 MiB versus 714.4 MiB of native history-slot capacity, approximately 25× smaller. On the evaluated tasks, recent cached content supports high success without retaining the full past; measuring that dependence requires controlling pool capacity separately from the attention window.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.