acceptodds
Under review as a conference paper at ICLR 2027

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

Abstract

Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between reasoning and trajectory generation. However, the robustness of such systems under input perturbations remains largely unexplored. We show that these models are highly vulnerable to input perturbations, achieving up to 100% attack success rate (ASR) on reasoning in open-loop and up to 72% on trajectory manipulation in closed-loop simulation, leading to increased collision rates and degraded safety metrics. Using NVIDIA's recent Alpamayo models as representative industry-developed VLAs, we conduct a systematic black-box study of reasoning-enabled driving VLAs and evaluate how perturbations affect reasoning and driving behavior. We introduce ReasonBreak, a reasoning-aware evaluation framework that captures semantic and structural changes in reasoning, trajectory degradation, reasoning–trajectory coupling, and downstream safety impact. Our analysis reveals that reasoning and trajectory changes are only weakly coupled under perturbation, with substantial changes in one output surface often occurring without corresponding changes in the other. We also introduce a benchmark for evaluating attacks and defenses on reasoning–trajectory interactions in autonomous driving. Our results highlight the need for rigorous evaluation and improved defenses to ensure the safety of reasoning-enabled VLA systems in autonomous driving. Implementation code and data are available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.