acceptodds
Under review as a conference paper at ICLR 2027

VLAs Don’t Always Do What They Say: Towards Measuring and Improving Alignment in Self-Driving VLAs

Abstract

While Vision-Language-Action (VLA) driving models generate textual reasoning traces alongside planned trajectories, these explanations often fail to align with actual execution. We address this misalignment between the reasoning and trajectory through a two part investigation: First, we conduct a mechanistic analysis on Alpamayo (a SoTA driving policy), applying causal interventions including cross-attention masking, prefix/suffix sweeps, clause reordering, logit-level alternatives, and paraphrasing. Our analysis reveals that reasoning traces are heavily front-loaded with maneuver commitments (where clause reordering induces larger trajectory errors than full reasoning masking) and are often superseded by superficial shortcuts in dense scenarios such as complex intersections and work zones. We synthesize these probing techniques into a standardized multi-axis faithfulness benchmark for VLA driving. Second, to prevent policies from exploiting naive alignment metrics, we propose TRACR (Trajectory-Reasoning Aligned Code Reward), a post-training framework using GRPO with a corpus of faithfully validated reward functions that leverage code subtask generation to enforce scene-grounding and logical consistency. Evaluated on held-out OOD scenarios from our benchmark, TRACR outperforms standard RL baselines (such as deterministic rules and LLM judges), improving reasoning-action consistency by over 17% and faithfulness by over 9% while preserving baseline trajectory accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.