PapoDrive: Pattern-Aware Policy Optimization for Safe and Efficient Driving Vision-Language-Action Model
Abstract
Reinforcement learning (RL)-based post-training has become a mainstream paradigm for improving the safety of vision-language-action (VLA) models in autonomous driving, yet the safety gains often come at the cost of overly conservative behaviors and significant efficiency degradation. To address this safety-efficiency trade-off, we propose PapoDrive, a Pattern-Aware Policy Optimization framework for the RL-based post-training. Specifically, we model the VLA planning as a two-stage decision, where the large language model (LLM) first selects a macro reasoning pattern, and the planner then decodes a trajectory conditioned on it. Under this view, the conventional KL-regularized RL essentially acts as a reweighting mechanism over the reference reasoning patterns, with the resulting efficiency drift governed by the covariance between the pattern-level efficiency and safety rewards, which is significantly negative in many driving scenarios and thus becomes the source of this trade-off. Guided with this insight, we develop a pattern-level step-size modulation (PSM) mechanism, allowing efficient patterns to move further from the reference policy while keeping inefficient ones close to it. We further design a lightweight pattern extraction scheme, such that PSM can be integrated into GRPO with a negligible overhead. PapoDrive improves the Success Rate by 5.72 and Efficiency by 6.15 over the ORION baseline on Bench2Drive benchmark, and achieves the best safety metrics while limiting efficiency drop to 2.2%, compared with 10.2% for GRPO on high-risk Navsim scenarios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.