acceptodds
Under review as a conference paper at ICLR 2027

CoSA-Drive: Counterfactual Safety Authorization for Vision-Language-Action Models in Autonomous Driving

Abstract

Vision‑language‑action (VLA) models have shown strong capabilities in scene understanding, task reasoning, and trajectory planning for autonomous driving. However, driving VLAs rely on a decision chain that combines visual observations, temporal context, vehicle states, and task instructions. Corruption of any source can propagate through this chain and induce unsafe behavior. Existing defenses mainly detect abnormal inputs or restore corrupted observations. Detection only indicates that an input may be unreliable, while restoration does not establish that the resulting driving behavior is safe to execute. We therefore reformulate VLA defense as counterfactual behavior authorization and introduce CoSA‑Drive. The framework first audits multimodal evidence and constructs heterogeneous behavioral candidates using specialized defense experts. Its core authorizer, CERA, independently evaluates each candidate using runtime evidence and predicts its safety, intervention‑induced harm, and task utility. Calibrated confidence bounds authorize only candidates that satisfy the risk constraint. Defense experts can propose behaviors but cannot execute them directly. A deterministic DriveContract further checks traffic rules, task permissions, and operational constraints, and triggers replanning or a minimal‑risk maneuver when no candidate is admissible. Experiments across four driving‑VLA backbones and reactive CARLA environments show that CoSA‑Drive recovers 62.2–95.5% of attack‑induced degradation and improves the safety–utility trade‑off under multimodal attacks and distribution shifts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.