From Knowing to Abstaining: Bridging the Representation–Action Gap in Vision–Language Models
Abstract
The ability of VLMs to correctly abstain from answering unanswerable questions is as important as their ability to generate accurate answers to answerable ones. Recently, several benchmarks have emerged to evaluate and improve VLM abstention. However, they suffer from substantial limitations. First, their samples often contain unintended shortcut cues in either the images or questions that inadvertently reveal answerability; moreover, they typically provide an explicit “unanswerable” option, preventing an accurate assessment of whether VLMs can abstain spontaneously. Second, when used as training data, they generally provide only binary labels without fine-grained explanations to support deeper supervision. To address these limitations, we introduce Visual Answerability Diagnosis with Rationales (VAD-R), a new benchmark constructed through a two-stage pipeline of shortcut filtering and quality verification to prevent answerability leakage. Each example is annotated with step-by-step rationales and causal evidence-gap labels. Evaluation of state-of-the-art open- and closed-source VLMs on VAD-R reveals strikingly limited spontaneous abstention capabilities, with the two model groups achieving average recall rates of only 11.4% and 16.3%, respectively. We then conduct probing analyses, revealing that hidden-state representations in certain layers can effectively distinguish answerability, yet this internal distinction fails to manifest in VLMs’ final responses. Motivated by this observation, we introduce Rep2Act, a representation-to-action alignment method that translates latent answerability awareness into explicit abstention decisions. Rep2Act substantially improves action accuracy on VAD-R from 56.67% to 86.33% for Qwen2.5-VL-3B and from 59.33% to 88.67% for Qwen2.5-VL-7B. On the out-of-distribution TUBench, Rep2Act achieves an average F1 score of 53.3% with only a 3B model, surpassing the closed-source GPT-4 Turbo and GPT-4o by 16.2% and 1.1%, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.