acceptodds
Under review as a conference paper at ICLR 2027

MM-ImpBench: Measuring Multimodal Implicitness for Graded Evaluation of Vision–Language Models

Abstract

Vision–language models (VLMs) perform well across diverse multimodal tasks, while their ability to recognize implied meaning remains questionable, which is important for safety-critical tasks such as detecting coded hate and toxic language. Existing benchmarks largely characterize implicitness through phenomenon-specific categories or binary labels, providing little indication of how indirectly an image–text pair communicates its meaning or how model performance changes as such messages become more indirect. To capture this missing dimension, we argue that multimodal implicitness should be treated as a continuous, task-independent property that can be measured directly, rather than reduced to a phenomenon-specific label, and introduce MM-ImpScore, a reference-free metric for multimodal implicitness, grounded in the discrepancy between an image–text pair's surface compositional meaning and the pragmatic meaning that emerges from their joint interpretation. We train the model on a dataset of controlled semantic families that vary communicative directness while approximately preserving the intended proposition. Building on this measure, we introduce MM-ImpBench, which scores samples from six existing datasets spanning hate speech, offensive content, sarcasm, and multimodal safety and stratifies them into distinct implicitness levels. Evaluating vision–language detectors reveals a recurring decline in detection accuracy as indirectness increases. Since difficulty concentrates in the high-implicitness regime, we use MM-ImpScore to route only such inputs to detailed reasoning, matching or improving always-on reasoning at lower latency. We will release our code and data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.