acceptodds
Under review as a conference paper at ICLR 2027

Discovering Interpretable Failure Modes of Vision Language Models

Abstract

Vision Language Models (VLMs) are increasingly deployed in safety-critical applications due to their general-purpose reasoning and adaptability with minimal domain-specific engineering. However, they fail repeatedly under specific, naturally occurring conditions, constituting failure modes. We present REVELIO, a framework for the systematic discovery of interpretable failure modes in VLMs. We formally define a failure mode as a combination of interpretable, domain-specific concepts such as proximity of a pedestrian or weather conditions, for which a target VLM fails consistently. Discovering them requires searching an exponentially large, discrete combinatorial space, which REVELIO casts as a black-box search over concept sets under compatibility constraints, adapting two search strategies to it: a diversity-aware beam search and Gaussian-Process-based Thompson Sampling for exploring larger concept combinations. Applying REVELIO to autonomous driving and indoor robotics surfaces failure modes of state-of-the-art VLMs that aggregate benchmarks do not localize. In driving scenarios, VLMs either miss n-lane obstacles, choosing actions that cause simulated collisions, or over-react to distant ones. In indoor settings, models either overlook hazards or exhibit overly cautious behaviors that trigger false alarms and degrade efficiency. By surfacing structured failure modes, REVELIO provides developers with actionable diagnoses to guide targeted VLM safety remediations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.