PROACTIVEDRIVE: FUTURE-AWARE ADAPTIVE THINKING FOR VISION-LANGUAGE-ACTION DRIVING
Abstract
Anticipating near-future scene evolution has shown strong potential for autonomous driving by enabling policies to foresee how their surroundings may change before acting. However, bringing this capability into vision-language-action (VLA) driving remains challenging, as future prediction and language reasoning can both increase inference cost in closed-loop deployment. We present ProActiveDrive, an adaptive-thinking framework for efficient VLA autonomous driving. Built on a SimLingo-style language-command policy, ProActiveDrive combines lightweight future-aware representation learning with selective reasoning at inference time. Instead of relying on dense generative rollouts, the model uses a compact feature-level future predictor to provide anticipatory context to the driving policy. At the same time, adaptive thinking avoids unnecessary language generation by activating explicit reasoning only when it is useful. On Bench2Drive-220, ProActiveDrive improves over the SimLingo language-command baseline, increasing the driving score from 86.08 to 88.20 and the success rate from 67.27% to 71.30%. Ablation studies show that adaptive reasoning improves inference efficiency and that compact near-future prediction is more effective than heavier multi-horizon prediction in closed-loop driving. Evaluation on Fail2Drive further highlights the remaining challenge of out-of-distribution generalization. Our results suggest that combining efficient reasoning with lightweight future-aware prediction can improve both efficiency and decision quality in embodied VLA systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.