acceptodds
Under review as a conference paper at ICLR 2027

TRUST-AD: Evaluating Textual Bias in MLLMs in Autonomous Driving Context

Abstract

Multimodal Large Language Models (MLLMs) are increasingly integrated into autonomous driving to handle complex driving situations that require more than basic object detection. In these tasks, models are expected to ground their reasoning in visual scene evidence, apply relevant traffic rules, and select safe driving actions. However, it remains an open question whether MLLMs actually perform this multi-step visual reasoning or instead depend on parametric knowledge of the underlying LLM and textual context to produce decisions. We introduce **TRUST-AD**, a benchmark designed to test whether driving decisions remain grounded in visual evidence or biased through textual descriptions of the scene. **TRUST-AD** comprises 14,238 test cases across six safety-critical driving categories. By keeping the visual scene fixed while systematically intervening on the textual prompt describing the scene, the benchmark measures decision stability, vulnerability to misleading textual cues, and error recovery. In addition, an intervention-based attribution framework identifies whether incorrect decisions stem from visual-grounding failures or rule-reasoning failures. Evaluating both general-purpose and driving-specific MLLMs shows that models frequently abandon visual observations to follow textual cues, even when baseline accuracy suggests strong performance. These findings indicate that high task accuracy can mask a reliance on language shortcuts rather than grounded visual reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.