Seeing Is Not Enough: Vision-Language Models Perceive Evidence but Fail to Act
Abstract
Vision-language models (VLMs) achieve strong performance on visual question answering benchmarks, yet often produce decisions that contradict visual evidence they have already correctly perceived. Existing work does not distinguish perceptual failure, where relevant evidence is not perceived, from process failure, where correctly perceived evidence fails to constrain the final decision. To address this gap, we first introduce VPAC-Bench, a benchmark of nine real-image process families in which each image is annotated with its current activity stage and the specific stage transition it falls near eg. whether a tool has made contact with an object (setup vs. active), or whether the intended outcome has been achieved (active vs. finished). Building on this structure, we develop SRT (State–Relevance–Target), a family of structured process-prior interventions that constrain how a model must use visible evidence before committing to a final answer. First, process failure is widespread and distinct from perceptual failure: models that correctly enumerate visual candidates still over-commit to a single answer at rates exceeding 95%, and an explicit process-structured intervention reduces this to below 13% without degrading unambiguous-referent performance. Second, process-prior transfer is model-sensitive: the same structured prior helps some models substantially while leaving others unchanged or degraded, and generic SRT alone does not reliably outperform strong CoT-style baselines. Third, when the correct stage transition is known, boundary-aligned SRT substantially outperforms generic process prompting and all CoT-style baselines across all four standardized process families (assembly, physical state transition, navigation/traffic, and object-use affordance). These findings indicate that the value of a process prior depends critically on its alignment with the image's specific decision boundary, motivating boundary-aware prior selection as a design strategy for process-grounded visual reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.