Decoupling Planning and Execution: A Cerebrum-Cerebellum Framework for Robust Visual-Language-Action Models
Abstract
Visual-Language-Action (VLA) models have demonstrated remarkable capabilities in end-to-end robotic manipulation by integrating multimodal information. However, these models frequently degenerate into static vision-action mappings, resulting in a critical failure to adhere to language instructions under complex constraints or target changes. We identify that this limitation stems from an inherent modality imbalance where physical actions correlate more strongly with visual states than with abstract linguistic descriptions. Consequently, models spontaneously develop a "vision-dominated bias" that marginalizes linguistic guidance during training. To address this, we propose a bio-inspired "Cerebrum-Cerebellum" architecture to structurally decouple task planning from execution. The language-dominant Cerebrum Module functions as a high-level planner to parse instructions into structured sub-tasks and tokens with visual information, while the vision-dominant Cerebellum Module executes these plans with real-time feedback. Across LIBERO-Plus and RoboTwin 2.0, our method improves instruction-following robustness, outperforming by 5.28 and 9.20 percentage points, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.