acceptodds
Under review as a conference paper at ICLR 2027

Decoupling Planning and Execution: A Cerebrum-Cerebellum Framework for Robust Visual-Language-Action Models

Abstract

Visual-Language-Action (VLA) models have demonstrated remarkable capabilities in end-to-end robotic manipulation by integrating multimodal information. However, these models frequently degenerate into static vision-action mappings, resulting in a critical failure to adhere to language instructions under complex constraints or target changes. We identify that this limitation stems from an inherent modality imbalance where physical actions correlate more strongly with visual states than with abstract linguistic descriptions. Consequently, models spontaneously develop a "vision-dominated bias" that marginalizes linguistic guidance during training. To address this, we propose a bio-inspired "Cerebrum-Cerebellum" architecture to structurally decouple task planning from execution. The language-dominant Cerebrum Module functions as a high-level planner to parse instructions into structured sub-tasks and tokens with visual information, while the vision-dominant Cerebellum Module executes these plans with real-time feedback. Across LIBERO-Plus and RoboTwin 2.0, our method improves instruction-following robustness, outperforming by 5.28 and 9.20 percentage points, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.