acceptodds
Under review as a conference paper at ICLR 2027

SimulVision: Diagnosing and Recovering Multi-Object Assignment in Vision-Language Models

Abstract

Multi-object errors in vision-language models (VLMs) are often interpreted as failures of visual binding or spatial reasoning, yet an incorrect global response does not reveal whether assignment-relevant information is genuinely unavailable. Inspired by the distinction between component access and multi-item integration, we study identity-location assignment under controlled evaluation. A component-preserving two-object diagnostic initially suggests a substantial mapping deficit: Qwen2.5-VL-7B fails on 31.4% of component-correct items under a fixed answer order. However, only 6.4% are incorrect under both answer orders, showing that many apparent assignment failures are response-unstable. We therefore introduce a closed-vocabulary multi-object assignment task that removes answer-position bias and free-form naming variability. On Qwen, local likelihood probing recovers substantial assignment-relevant signal even when monolithic response accuracy is low. Across both Qwen2.5-VL-7B and InternVL3.5-4B, enforcing one-to-one structure on the same local compatibility scores substantially improves exact assignment accuracy. We instantiate this intervention as Factorized Global Assignment (FGA), a training-free structured readout, and replicate the effect on an independently frozen 300-image confirmation set. These results show that multi-object evaluation should distinguish what assignment evidence can be elicited locally from how reliably several assignments are assembled into a globally consistent response.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.