Aesthetic Visual Reasoning with Multiple Images
Abstract
Aesthetic evaluation in photography often depends on collections rather than single images, as a photographer’s style emerges through recurring patterns and themes across multiple works. This motivates moving beyond single-image assessment toward group-level aesthetic modeling, which better captures holistic characteristics and supports applications such as photo recommendation, education, and style transfer. Existing Multimodal Large Language Models (MLLMs), however, largely focus on single-image or pairwise comparisons, which limits their group-level reasoning due to two key limitations: the absence of datasets with group-level aesthetic annotations and insufficient multi-image understanding in current models. To address these gaps, we introduce the first dataset for group-level aesthetic reasoning in photography, analyze the behavior of MLLMs in multi-image contexts, and uncover some unique properties of aesthetic reasoning across image groups. Building on these insights, we fine-tune a multi-image MLLM and introduce a lightweight visual triager that re-orders images based on their perceived difficulty, addressing the model’s sensitivity to image order and improving inference on challenging images. Experiments demonstrate that our approach significantly outperforms existing models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.