acceptodds
Under review as a conference paper at ICLR 2027

ReTranDrive: A Reinforced Transition-Aware Framework for Vision-Language-Action Autonomous Driving

Abstract

Autonomous driving policies are commonly optimized by evaluating the quality of each predicted plan in isolation, despite being executed through continuous replanning. This formulation overlooks a distinct aspect of sequential decision-making: the quality of transitions between consecutive plans. As a result, a planner can produce individually high-quality trajectories while exhibiting unstable behavior over time, repeatedly altering its intended motion without a corresponding change in the scene. The missing ingredient is an explicit objective over inter-plan transitions, capturing whether successive predictions evolve coherently with the changing scene rather than merely optimizing each plan in isolation. We present ReTranDrive, which defines this objective by shifting the optimization target from single-step planning quality to sequential execution coherence: scoring a plan against the current observation alone becomes a transition-aware formulation over consecutive plan pairs. Applied to a VLM-conditioned diffusion planner as a differentiable term during imitation and a simulator-verifiable reward during reinforcement, it regularizes only those inter-plan changes not warranted by scene evolution rather than enforcing unconditional temporal consistency. To resolve the credit-assignment challenge inherent in pairwise scoring, we introduce an asymmetric scheme that assigns inconsistency strictly to the decision responsible for it. On the NAVSIM benchmark, ReTranDrive attains a state-of-the-art EPDMS of 90.94 under the extended metric and a PDMS of 94.40, reaching 99.58% of human driving performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.