acceptodds
Under review as a conference paper at ICLR 2027

SUBTEXT: Can Image Generators Infer What Captions Leave Implicit?

Abstract

Can image generators infer an outcome that a caption implies but does not state? Captions routinely omit direct descriptions of outcomes and rely on the reader to infer them from context. For example, a caption describing a chocolate bar left on a dashboard on a hot afternoon implies that the bar has melted. We introduce SUBTEXT, a benchmark of 140 minimal image pairs that differ in one physical or socio-emotional outcome within an otherwise matched scene. Each image has an implicit caption that specifies only the antecedent context and an explicit caption that is identical except for an inserted outcome term such as "melted". Following work on diffusion classifiers, we use seven frozen image generators as zero-shot classifiers that match captions to images and images to captions. We compare them with three contrastive models, Claude Fable 5.1, and human participants. With implicit captions, the generators select the correct image at near-chance rates and fall below chance when both captions of a pair must be matched. Human participants and Claude Fable 5.1 select the correct image in at least 93% of cases. Explicit captions significantly improve accuracy for six of the seven generators, which indicates that inference adds difficulty beyond visual discrimination. Unlike prior results on compositional benchmarks, the generators show no advantage over contrastive models, even when the two share a text encoder. These results suggest that learning to generate realistic images does not by itself teach a model to reliably link context to visible outcomes. SUBTEXT therefore provides a controlled test of whether future generators close the gap between what a caption states and what it implies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.