VarioDrive: Adaptive and Diverse Reasoning for End-to-End Autonomous Driving
Abstract
Vision-language-action (VLA) models have shown a promising effectiveness in autonomous driving planning. Existing approaches, however, primarily optimize the final trajectory, providing limited feedback on whether the intermediate reasoning is grounded in the scene and actually supports the plan. Moreover, they typically commit to a single reasoning path per scene, leaving alternative behaviors with different safety, efficiency, and comfort trade-offs unexplored. To address these issues, in this paper, we propose VarioDrive, a planning-oriented post-training framework that aligns structured reasoning with trajectory generation and enables adaptive, preference-conditioned planning within a single autoregressive model. Specifically, we first introduce a time-indexed waypoint tokenization that factorizes each future pose into time-indexed position and heading tokens, together with a geometry-aware soft supervision that makes token prediction aware of spatial distance, so that trajectories can be directly decoded without an external planner. Then, we organize each response into a scene description, driving-intent reasoning, and a planning answer, and design stage-decomposed rewards that separately evaluate visual grounding, intent-trajectory alignment, and preference-dependent trajectory quality during RL post-training. We further develop an adaptive fast-slow inference mechanism that generates a single plan for simple scenes while generating and comparing safety-, efficiency-, and comfort-oriented plans for complex scenes before selecting the final trajectory. Experiments on NAVSIM, nuScenes, and Bench2Drive demonstrate a competitive planning performance. On NAVSIM, stage-decomposed rewards improve PDMS by 2.09 points over answer-only RL, and the resulting model reaches 92.42 PDMS with multi-preference selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.