acceptodds
Under review as a conference paper at ICLR 2027

Presence Is Not Faithfulness: Figurative Vehicle Intrusion in Text-to-Image Generation

Abstract

Text-to-image (TTI) models increasingly generate high-quality images from natural-language prompts, yet figurative language exposes a failure: a vehicle that should guide the depiction of a tenor may instead be rendered as a visible object. We call this failure **Figurative Vehicle Intrusion**: the intruding content is textually licensed, but it is assigned the wrong visual role, showing that visual presence is not always faithfulness and that presence-oriented evaluation can miss such errors. To study it systematically, we introduce **V**ehicle **I**ntrusion and **S**emantic **T**enor **A**ssessment (**VISTA**), a multilingual benchmark of figurative prompts organized by Figurative Form and Mapping Mechanism. We further propose **V-Score**, a diagnostic question-answering metric that evaluates role-aware figurative faithfulness in generated images. Evaluations on recent high-performing TTI models show that vehicle intrusion persists across languages and figurative categories. As a lightweight mitigation, we introduce **VISTA-Guard**, which partially reduces vehicle intrusion and suggests a practical path toward more figuratively faithful TTI generation. All resources will be released publicly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.