acceptodds
Under review as a conference paper at ICLR 2027

: Evaluating Visual Simile Canonicity in Text-to-Image Models

Abstract

When a text-to-image (T2I) model is asked for "a curved moon", it will draw a crescent with ease. However, when it is prompted with a simile such as "a moon that looks like a boat", it often draws a plain boat, a moon beside a boat, a full moon, or a moon with apparent masts and sails. Such difficulty is because *"looks like"*'s implication shifts with every pair of things being compared to in the prompt, e.g, a moon that resembles a boat takes a boat's curve, while a cherry that resembles a ruby takes a ruby's color. To bridge this gap so that T2I models can better handle such prompts with implicated resemblance, we thus formalize the **Visual Simile Generation (VSG)** task: given a simile in form of " that look(s) like ", a T2I model must: 1., install as the dominant subject, 2., identify and transfer the implicated attribute from the comparison object onto the subject , and 3., render no concrete features of literally. We curate **VSG-Bench**, consisting of 549 simile prompts spanning 13 resemblance relations over four axes. At evaluation, each prompt licenses the one attribute of a generation is expected to transfer, and names the concrete features of it must leave out. We also establish the composite CanonScore metric, which scores a generation on each of the three requirements separately: for the subject, for the transferred attribute, and for staying clear of 's concrete features. We benchmark 16 recent T2I systems spanning various architectures. We find no current T2I system can satisfy all three requirements simultaneously. More signifcantly, corroborated by human raters, our full paradigm effectively exposes **the resemblance-literalization tradeoff**, that the T2I systems conveying the intended resemblance attribute best are also the ones literalizing the most of the comparison objects.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.