Continuous Reasoning for Vision-Language-Action Models
Abstract
A robot's intermediate thoughts can be trained through the actions they enable, and a useful thought should serve more than the network that produced it. We introduce Continuous Reasoning (CR), a vision-language-action (VLA) framework that learns a sequence of continuous thought vectors directly from action supervision, without annotated reasoning traces. Our key idea is self-verification: a thought should support action prediction by both the policy that generates it and an evolving exponential-moving-average (EMA) consumer. The consumer's action-prediction error backpropagates to the code generator, training thoughts through their use for action prediction. Built on π0.5, CR generates thoughts sequentially through a Gaussian-prior bottleneck and uses the resulting prefix to condition block-causal action prediction. The EMA consumer is used only during training, so inference requires no separate verification pass. On LIBERO-PRO, CR improves mean success from 58.0% to 64.0% over its backbone and exceeds DeepThinkVLA-RL in all four suite means, with a 14.0-percentage-point overall advantage. Gains concentrate on position and task perturbations, where policies must adapt to changed layouts and goals. On TX-G2 and HSR, mean subtask success improves by 19.8 and 15.6 percentage points, respectively, with gains also extending to full-task completion. Every component ablation loses on those same two families, and freezing the consumer removes the benefit of self-verification entirely. These results establish action-supervised continuous reasoning as an effective interface for robust VLA control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.