When Should the Decision Change? Counterfactual Probing of Vision-Language Models for Autonomous Driving
Abstract
In autonomous driving, a small change in another road user's behavior can make a previously appropriate ego decision unsafe. How reliably do vision-language models adapt their driving decisions to such changes? We introduce DriveWHEN, a counterfactual probing framework that examines model decisions across controlled alternatives to the same traffic scene. The framework combines scene interventions, context-dependent action requirements, and paired evaluation to assess both responsiveness when a different action is required and stability when the original action remains valid. Grounded in real nuScenes scenes, we construct paired videos through four atomic interventions on existing scene entities: leftward and rightward lateral intrusions designed to introduce hazards, forward departures designed to alleviate hazards, and appearance replacements serving as controls. Human review independently assigns the required ego action to each original and counterfactual video, allowing the same motion intervention to warrant either retaining or revising the original action across different traffic contexts. Evaluation of five vision-language models reveals frequent failures to revise their actions when the altered traffic situation requires a different response. These findings expose a gap between producing a decision for an individual scene and responding to changes in the interactions that inform that decision. DriveWHEN tests whether models can distinguish visually similar traffic situations and revise their decisions only when the altered context warrants it, providing a targeted diagnostic of response reliability in safety-critical driving.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.