acceptodds
Under review as a conference paper at ICLR 2027

The Scene, Not the Shot: Measuring Inferred Composition in Text-to-Image Models

Abstract

Photographers stage emotion geometrically: awe by dwarfing a figure against a vast scene, dominance by tilting the camera upward. These conventions are rarely spelled out—no photographer is told to shrink the subject—yet they are encoded in the photographs text-to-image models are trained on. Models follow explicit framing instructions well; what they do by default, when the framing is not specified, has not been measured. We ask whether models infer the composition an intent calls for, through a benchmark of six axes—awe, romance, social group, power, vulnerability, and relative scale—whose prompts name the intent but never how to frame it. Each axis extracts a geometric quantity from the image and scores it against human photographic practice or an established convention, with bootstrap uncertainty throughout. The reference is a range of human practice, not a single correct answer or a recipe for the emotion. Across twelve recent models we find one consistent split, holding within every model: models get the scene right but not the camera. What lies within the scene—the depth of an awe landscape, the spacing of couples and groups—stays near human practice. How the camera captures it does not: the awe subject is framed too large or too uniformly, couples are shot from too far, and although power calls for a low camera angle and vulnerability a high one, every model keeps the camera at eye level in at least seven of ten generations, tilting it barely more often than when no intent is named. Relative scale compresses: small things render too large, vast things too small. No model dominates across axes. Models stage the scene; they do not shoot it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.