acceptodds
Under review as a conference paper at ICLR 2027

Aesthetic Visual Reasoning with Multiple Images

Abstract

Aesthetic evaluation in photography often depends on collections rather than single images, as a photographer’s style emerges through recurring patterns and themes across multiple works. This motivates moving beyond single-image assessment toward group-level aesthetic modeling, which better captures holistic characteristics and supports applications such as photo recommendation, education, and style transfer. Existing Multimodal Large Language Models (MLLMs), however, largely focus on single-image or pairwise comparisons, which limits their group-level reasoning due to two key limitations: the absence of datasets with group-level aesthetic annotations and insufficient multi-image understanding in current models. To address these gaps, we introduce the first dataset for group-level aesthetic reasoning in photography, analyze the behavior of MLLMs in multi-image contexts, and uncover some unique properties of aesthetic reasoning across image groups. Building on these insights, we fine-tune a multi-image MLLM and introduce a lightweight visual triager that re-orders images based on their perceived difficulty, addressing the model’s sensitivity to image order and improving inference on challenging images. Experiments demonstrate that our approach significantly outperforms existing models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.