acceptodds
Under review as a conference paper at ICLR 2027

Less Is an Act: Probing Selective Judgment in Vision-Language Models through Image Translation

Abstract

In-image machine translation (IIMT) has advanced rapidly. Existing systems and benchmarks are largely concerned with two things: the visual quality of the rendered output and the fidelity of the translation itself. However, we argue these objectives sit closer to a design-and-application problem than to a core question about model capability: when a user simply asks to “translate this image,” does the model understand what should be translated and what left untouched? Brand names, ingredient codes, and text physically part of an object must be preserved for translation to serve the user. We find that state-of-the-art vision-language models (VLMs), given no guidance, systematically mistranslate what should be left untouched, most often by over-translating: rewriting brand names and other text that competent translators would preserve. We introduce selective translation, a new task that shifts the focus from translation quality to selective judgment and restraint: deciding, for each piece of text, whether to translate or preserve it. We take a new approach to in-image translation, building SPOT-Bench, a benchmark that evaluates what a model chooses to translate rather than how well it translates, with constructional ground truth, derived from how each image is made rather than annotated by taste, across advertisements, news, and captioned video. Evauating current proprietary and open VLMs, we find a large and consistent gap between what models know the rules to be and what they do: they fail not from missing knowledge but from failing to apply it when acting on the image. Our analysis shows that making the translate/preserve decision explicit substantially narrows this gap, pointing toward more selective, context-aware VLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.