acceptodds
Under review as a conference paper at ICLR 2027

NextShot: Learning Source-Grounded Structured Composition Targets for Portrait Photography

Abstract

Modern visual models increasingly move beyond recognizing isolated content toward structured understanding of scene entities, semantics, and relationships. Relational information is useful not only for characterizing what is observed: when a task requires inferring a desired configuration from existing visual content, that configuration often depends on interactions among entities and their global context. We study this problem in photographic composition, where relevant scene elements are already present and composition depends on how people and environmental regions are positioned and scaled relative to one another and the frame. We formulate Region-Level Composition Target Prediction, which preserves source-region identities while jointly predicting regional geometry, target frame aspect ratio, and person framing. Because photographic demonstrations typically provide only whole-image source–reference outcomes, direct region-level supervision is unavailable. We recover structured supervision by transferring reference geometry to source identities through partial correspondence while preserving each reference as a complete alternative arrangement, complemented by single-image aesthetic annotations. Our predictor, NextShot, models region–region and region–frame relations, integrates multimodal scene semantics, and conditions regional geometry on predicted frame and person framing. NextShot improves centroid and visible-area prediction over an autoregressive alternative. Two independent large multimodal evaluators from different model families also consistently prefer its complete layouts over alternative predictors and key model variants. In a 100-case guided-recapture study, five independent human evaluators prefer photographs taken with NextShot guidance in 78.0% of 500 judgments. These results establish source-grounded structured targets as a learnable representation for explicit spatial goal prediction from relational scene information.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.