RLA-Net: Rubric-Grounded Learning for Decomposable Image Aesthetic Assessment
Abstract
A single aesthetic score summarizes preference but does not explain the visual strengths, weaknesses, and trade-offs behind it. Learning interpretable image aesthetic assessment (IAA) from aggregate ratings is challenging: the ratings leave individual aesthetic states underdetermined, prompt–feature similarity does not ensure attribute-specific visual responses, and rubric states alone do not specify their image-dependent importance. We introduce RLA-Net, a CLIP-based framework that jointly estimates readable aesthetic states and learns their adaptive scoring roles through rubric-grounded score composition. Local rubric grounding estimates grades by matching visual evidence to three language-defined reference states for each of eight aesthetic dimensions. Contrastive Rubric Alignment (CRA) complements overall-score supervision with source-relative, rubric-targeted synthetic pairs that constrain global image–anchor responses in prescribed directions. Gated Aesthetic Scoring combines rubric grades, visual descriptors, and image context into explicit prediction terms, allowing similar states to play different scoring roles across images. Training uses image-level ratings, fixed rubric anchors, and synthetic pairs, without per-image rubric annotations, comments, critiques, or teacher-generated targets. Experiments on AVA, PARA, TAD66K, and FLICKR-AES demonstrate competitive performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.