COCOA: Action-Conditioned Correction for Vision-Language-Action Policies
Abstract
Vision-language-action (VLA) models have shown strong performance on complex robot manipulation tasks by generating low-level action chunks from multimodal observations and language instructions. However, an erroneous action chunk can derail long-horizon execution and lead to task failure. To address this challenge, we propose , an action-conditioned correction training framework that learns a separate corrective policy to revise action chunks proposed by a base VLA before execution. Given the current visual observations, language instruction, robot state, and a proposed action chunk, the corrective policy generates a corrected action chunk for execution. To train the corrective policy, introduces a two-stage training procedure that combines supervised flow-matching training with correction preference optimization for flow-based VLAs. We apply to three flow-based VLA backbones, including , FLOWER VLA, and Xiaomi-Robotics-0. Our experiments show that achieves consistent improvements on two simulation benchmarks, SimplerEnv and CALVIN, as well as on real-robot tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.