acceptodds
Under review as a conference paper at ICLR 2027

LongContextVLA: Reasoning over Hundreds of Images for Robot Control

Abstract

Vision-language-action (VLA) models typically condition on only the current observation or a small set of past keyframes. This limits their ability to solve long-horizon tasks that require persistent visual memory. We develop a VLA that conditions on more than 160 images of visual history, represented with an average of 72 tokens per image, while maintaining real-time inference speed. Our approach builds on Qwen3.5-4B, a multimodal model with a 256K-token context window, which we fine-tune for robotic control. We find that long visual context alone is insufficient for robust generalization, and train the VLA to reason over and narrate events in its visual history, using parallel decoding to avoid the latency cost of autoregressive reasoning. By caching past visual tokens and processing only newly arriving images, our streaming implementation achieves per-step latency close to that of a single-image VLA. We evaluate our approach on a long-horizon shell game and object retrieval task, achieving over 80% success in unseen generalization scenarios, as well as competitive performance on the RoboMME benchmark. Finally, we demonstrate memory-based prompting, where the robot observes a human demonstration and reproduces the demonstrated behavior, with the ability to generalize to unseen scenarios. Together, our results show that scalable visual memory, learned reasoning, and streaming inference enable real-time, long-horizon robot control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.