CABLE: Evaluating Selective Low-Level Visual Evidence Tracking in MLLMs
Abstract
Multimodal large language models (MLLMs) are increasingly applied to image quality assessment, degradation perception, and other low-level visual tasks. Yet isolated image-question accuracy cannot distinguish responses grounded in query-relevant visual evidence from those supported by linguistic and high-level semantic cues. We formulate this problem as selective evidence tracking: a model should revise its response when relevant evidence changes the correct result and preserve it when the result remains unchanged. To evaluate this behavior, we introduce Controlled Assessment of Bundled Low-Level Evidence (CABLE), which organizes anchor instances and controlled variants into intervention-induced transitions. Conditioned on a correct parent response, CABLE measures response adaptation (RA) on altered transitions and response preservation (RP) on invariant transitions, while decomposing adaptation failures into response inertia (RI) and misdirected updates (MU). Across 29 general-purpose and low-level-vision-specialized MLLMs, models preserve correct judgments under invariant changes far more reliably than they revise them when visual evidence changes the required result. More importantly, response revision is substantially weaker for visual than query-induced changes, with adaptation failures dominated by response inertia rather than misdirected updates. Fine-grained analyses further reveal systematic effects of visual relevance, query grounding, and linguistic elicitation that isolated accuracy obscures. CABLE therefore complements pointwise accuracy with transition-level diagnostics of whether model responses remain responsive to the low-level visual evidence that should determine them.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.