acceptodds
Under review as a conference paper at ICLR 2027

MindDrive: A Map-Informed Adaptive Decision Method for Efficient End-to-End Driving

Abstract

Recent advances in Vision-Language-Action (VLA) models have demonstrated the potential of combining multimodal understanding, reasoning, and trajectory generation for end-to-end autonomous driving. However, existing methods often infer traffic context directly from multi-view images, making them susceptible to hallucinations about planning-critical scene elements. Moreover, their reasoning strategies are typically insensitive to scene complexity, leading to insufficient deliberation in challenging scenarios and unnecessary reasoning overhead in routine ones. In this paper, we propose MindDrive, a ap-formed Adaptive ecision VLA model for efficient driving. First, we design a language-map construction and retrieval pipeline that converts structured map elements into spatially indexed natural-language annotations, providing accurate, decision-relevant scene priors.Then, during supervised fine-tuning stage, the model learns to assess scenario complexity and accordingly perform either direct action generation or deliberative reasoning. Finally, we introduce a reinforcement learning-based post-training strategy that contrasts different decision modes within the same scene, further improving planning performance and reasoning efficiency. Extensive experiments on NAVSIM and nuScenes benchmarks demonstrate the competitive performance of MindDrive.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.