Looping Visual Tokens for Scaling Multimodal Intelligence
Abstract
We introduce Looping Visual Tokens (LVT), a mechanism for learning visual and spatial understanding through recurrent computation over visual tokens. LVT uses a shared transformer block to iteratively update a persistent visual state at the native image-token positions of a vision-language model (VLM). By reusing the same parameters and token positions across iterations, the block deepens visual computation without expanding the token sequence. We train this recurrent computation by aligning successive states with the backbone's perceptual representations of intermediate observations, and pass the final state to the model's task readout. Human videos and continuous environment-state sequences from simulation can provide supervision at scale, connecting LVT to learning from visual experience. We evaluate LVT on visual and spatial reasoning tasks. On fresh VSP maps beyond training path lengths, ordered supervision achieves 97.0% closed-loop success, compared with 93.2% for final-state supervision and 64.3% for answer-only recurrence. For unseen rotations at prescribed horizons, ordered supervision improves accuracy from 79.1% to 92.1%, while visual supervision alone reaches 87.2%. Textual process supervision remains stronger on the training grid size, while LVT generalizes better to larger grids. Decoding and interventions examine visual structure and state use in prediction. These results motivate LVT as a scalable mechanism for general-purpose visual learning. Anonymous repository: https://anonymous.4open.science/r/LoopVLM-6F55/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.