VLA-CoRe: Empowering Vision Language Action Model with Cognitive Reasoning
Abstract
Vision Language Action (VLA) models have demonstrated strong behavioral cloning performance on short-horizon manipulation tasks, yet they fundamentally lack cognitive awareness: the ability to estimate their own uncertainty, identify the source of potential failure, and selectively invoke specialized reasoning modules before committing to an action. We propose VLA-CoRe, a cognitive VLA framework that augments a pretrained multimodal LLM with an Uncertainty Monitor that gates three differentiable reasoning modules: (1) Perception Query, which resolves ambiguity by re-examining occluded or visually complex scene regions; (2) World Imagination, which performs counterfactual "what-if" simulations to predict the outcome of an action before physical execution; and (3) Causal Tracing, which attributes unexpected failures to specific past reasoning steps to revise the current plan, based on uncertainty estimates. The reasoning outcomes (from these three reasoning modules) are combined into a visual plan latent that conditions a Diffusion Transformer action decoder for robust low-level control. To train VLA-CoRe without expensive manual annotation, we introduce CIDAS (Cognitive Instruction Data for Autonomous robot reaSoning), a dataset of 170k conversation-style cognitive reasoning samples extracted from existing behavior cloning trajectories across seven structured task types. Training combines (i) a Supervised Fine-Tuning stage on CIDAS and (ii) Group Relative Policy Optimization (GRPO) guided by a cognitive objective. Crucially, our framework improves failure recovery rates by 21% and reduces expected calibration error by 38%, marking a significant step toward resilient & self-aware robots. Code and data will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.