acceptodds
Under review as a conference paper at ICLR 2027

CHIAA-Bench: Aligning Cinematic Human Image Aesthetic Assessment with Human Aesthetic Judgments

Abstract

Image aesthetic assessment has been widely studied for generic images, but remains underexplored for cinematic human images, whose aesthetic quality is mainly determined by human-centered cues such as facial clarity, expression, and atmosphere. To address this gap, we introduce a new task, **Cinematic Human Image Aesthetic Assessment (CHIAA)**, which aims to evaluate the aesthetic quality of cinematic human images from movies and TV series. We further present **CHIAA-Bench**, a new benchmark constructed from movie and TV frames as well as open-source cinematic datasets through human image filtering and a two-stage annotation pipeline with model scoring and expert refinement. Our annotation protocol covers four aesthetic dimensions and ten fine-grained sub-dimensions, producing interpretable five-level labels for cinematic human image aesthetics. Based on this benchmark, we develop **CineAesNet**, a lightweight model for efficient cinematic aesthetic assessment. Experiments show that CineAesNet consistently outperforms existing open-source aesthetic models and strong multimodal large language models on CHIAA-Bench while using substantially fewer parameters. In addition, our analysis shows that model-generated scores are useful for candidate filtering and low-resource warmup, whereas expert supervision remains essential for reliable learning of cinematic human aesthetics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.