Deconfounding and Compressing History Memory in Streaming Vision-Language-Action Policies
Abstract
How much explicit history does a streaming vision-language-action (VLA) policy need? A common test shortens its attention window, but this intervention can also shrink the physical key/value (KV) pool and change which history remains visible during denoising. We separate these quantities in LingBot-VA and show that an apparent collapse from 97.5% to 0% success on LIBERO-Long is a cache-capacity confound: prediction slots displace the previous chunk’s video tokens. With the window and pool controlled, retaining the immediately preceding chunk achieves 99% success versus 97% with full history; a three-task RoboTwin replication shows the same zero-to-one retention transition. Content interventions test what the retained chunk contributes. We then combine short retention with KV quantization. On LIBERO-Long, two chunks at K4/V2 achieve 99% success, with a paired difference of 0 percentage points (95% CI: [−3, 3]) from the native W = 30 FP reference. Their history-KV storage equivalent under simulated quantization is 28.5 MiB versus 714.4 MiB of native history-slot capacity, approximately 25× smaller. On the evaluated tasks, recent cached content supports high success without retaining the full past; measuring that dependence requires controlling pool capacity separately from the attention window.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.