The Action Doesn't End the Thought: Continuing Computation in Vision-Language-Action Models
Abstract
Most vision-language-action (VLA) models restart their action expert at every de- cision, discarding evidence the robot can no longer see and computation the next decision could reuse. We propose continuing computation: the expert carries its internal state across decisions and refines it with each new observation. We im- plement this in LycoReco, a VLA with a Continuous Thought Machine expert, and compare carry and reset experts trained identically except for state lifetime. When earlier cues differ but current inputs are identical, the carried state lets the robot act on information it no longer sees. It also saves expert computation: when replanning at every step on LIBERO, carry with one internal update per obser- vation outperforms reset with twelve. Carry’s successive action chunks are more consistent when nothing new is revealed, yet adapt when a remembered target moves. LycoReco is competitive with published VLAs on LIBERO (97.5%) and SIMPLER-WidowX (64.6%) and leads the compared systems on the MIKASA memory tasks (62.6%) without a separate memory module. State lifetime is there- fore a consequential architectural choice, and resetting the action expert at every decision should not be the default.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.