CoAMA: A Hierarchical Benchmark for Diagnosing Cross-Object Attribute Misbinding in Vision-Language Models
Abstract
Large vision-language models (LVLMs) are increasingly used for visual understanding and decision-making, yet they may still produce plausible but incorrect visual interpretations, a phenomenon known as hallucination. In multi-object scenes, such failures can occur when an attribute belonging to one object is incorrectly assigned to another, undermining reliable scene understanding and downstream reasoning. However, existing hallucination benchmarks mainly evaluate object existence, class recognition, or attribute presence, while largely overlooking whether perceived attributes are correctly bound to their corresponding object instances. To address this gap, we introduce Cross-object Attribute Misbinding Assessment (CoAMA), a hierarchical diagnostic benchmark for evaluating instance-level attribute binding through three progressive levels: multi-object instance recognition, target-object attribute recognition, and cross-object attribute binding verification. Its core design employs paired in-image attribute-swap probes, where each negative probe transfers a ground-truth attribute from one marked object to another, directly testing whether the model preserves instance-level attribute ownership when multiple attributes are visually present in the same scene. CoAMA contains 787 multi-object images, 3,935 marked instances, and 19,460 scorable attributes, with standardized annotation and scoring protocols for consistent and reproducible evaluation. We evaluate representative LVLMs and find that even the best-performing model, Qwen3.7-Plus, achieves only 67.69% PairAcc under L3-Oracle-Direct, while most models exhibit higher false-acceptance rates for attributes belonging to other marked objects than for attributes absent from the marked-object set. These findings reveal cross-object attribute misbinding as a critical reliability challenge and provide a controlled foundation for improving LVLM reliability in complex visual environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.