CausalDrive: Breaking Motion Inertia in VLA-based Autonomous Driving by Aligned Semantic Intervention
Abstract
As End-to-End autonomous driving advances into the Vision-Language-Action (VLA) era, models increasingly rely on natural language to generate human-aligned control. Despite this progress, a critical vulnerability remains: inertial driving, defined as the tendency to maintain historical motion patterns even when facing sudden hazards. We identify three factors underlying this issue: memory bottlenecks in temporal fusion, inertial bias in implicit temporal modeling, and decoupling of reasoning and action, which together allow misleading temporal inertia to dominate decision-making. To address this issue, we introduce CausalDrive, a novel VLA framework that explicitly bridges semantic reasoning and temporal action generation through language-guided causal intervention. Concretely, our architecture features three core innovations: 1) Streaming Sparse Temporal Fusion (SSTF) for maintaining efficient, long-term context via lightweight memory slots; 2) Task-Query-Guided Temporal Modeling to generate kinematically consistent ego-trajectories; and 3) Causal-Intervention Attention (CIA), a mechanism that dynamically overrides historical momentum when triggered by safety-critical language cues (e.g., “ignore momentum”). We further optimize CausalDrive through supervised fine-tuning and constrained Group Relative Policy Optimization (GRPO) guided by learned Reward/Cost Models. Extensive experiments demonstrate its effectiveness, achieving 87.41% accuracy on DriveAction, 0.30m trajectory error on DriveLM, and 0.0142 steering MAE on BDDA. Qualitative results further show its ability to suppress inertial bias during critical events, paving the way for highly reliable and interpretable autonomous driving systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.