ActionCritic: Action-Aware Self-Criticism for Reliable Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models provide a scalable interface for mapping multimodal observations and language instructions to robot actions. However, most VLA policies ex- ecute predicted actions without explicitly assessing whether those actions remain appropriate under the current scene con- figuration, where a locally incorrect command can cause an irreversible task failure. In this paper, we introduce Ac- tionCritic, a unified VLA framework that internalizes ac- tion verification, corrective instruction generation, and dy- namic re-planning within a single shared multimodal repre- sentation. Rather than relying on computationally expensive external simulators or disjointed auxiliary correction mod- ules, ActionCritic proactively evaluates the physical suitabil- ity of proposed event-triggered key actions against the ob- servation history and task semantics prior to execution. Un- suitable action proposals are preempted and translated into concise atomic corrections, which directly condition subse- quent action generation to enable bounded, local recovery. Training follows a two-stage schedule that first establishes critic capability with QA supervision and then mixes plan- ning and correction-conditioned re-planning while retaining a low-weight QA objective. Furthermore, ActionCritic learns a robust, deployment-aware decision boundary by compre- hensively leveraging both successful and failed trajectories while rigorously excluding failed motor commands from im- itation targets. Extensive evaluations across the SIMPLER and RoboTwin benchmarks demonstrate that ActionCritic sig- nificantly improves closed-loop task success, critic discrim- ination, and recovery effectiveness while maintaining strict inference efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.