acceptodds
Under review as a conference paper at ICLR 2027

Grounded Evaluation of Reasoning-Action Faithfulness in Vision-Language-Action Driving Models

Abstract

We study whether the chain-of-causation reasoning of vision-language-action (VLA) driving models is faithful to the actions they execute. Current evaluation relies on self-consistency checks and plausibility judges, which reward reasoning that reads coherently regardless of the trajectory the model drives. We instead evaluate faithfulness against two grounded references, the human-driven trajectory and human-verified expert reasoning, on Alpamayo-R1 and Alpamayo-1.5 across out-of-distribution driving events. Our main finding is that faithfulness is not plausibility: (i) both models plan accurately with average displacement error of 2.22 m and 3.68 m, yet their reasoning is fully faithful in only 39.6% and 33.1% of scenes, while reference-free signals rate 58-63% as faithful; (ii) when we hold reasoning fixed and progressively corrupt the trajectory, trajectory-sensitive metrics degrade monotonically, while trajectory-blind metrics, including the plausibility judge, stay constant; (iii) a validated counterfactual probe finds overt post hoc rationalization in fewer than 4% of scenes, so the gap reflects weak grounding rather than deception; and (iv) extending the evaluation to AutoVLA, which requires assistant prefill to produce reasoning, we find it less self-consistent than either Alpamayo model. We formalize grounded scorers, plausibility scorers and reasoning-action divergence, and prove that plausibility scorers are invariant to trajectory corruption by construction. Chain-of-causation explanations are therefore not yet a reliable basis for oversight of driving VLAs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.