Preferences Order, Ratings Anchor: From Fused Expert Aesthetic Ground Truth to Self-Distillation
Abstract
Pairwise preferences and pointwise ratings are the two dominant annotation protocols in image aesthetic assessment (IAA), yet existing benchmarks adopt only one, leaving their complementarity unmeasured on the same images and dimensions. We propose an approach to expert aesthetic evaluation based on anchored preference-rating fusion. To study this approach, we construct PPAINT, a benchmark of 150 Chinese paintings spanning five aesthetic dimensions. It contains matched ratings and 45,900 pairwise judgments collected from 15 domain experts. The matched design reveals complementary strengths: preferences yield more consistent ordinal rankings, while ratings anchor scores to the expert rating scale. Our approach combines these signals into fused expert ground truth using two distinct preference-to-score constructions, which produce nearly identical scores under shared anchor constraints. Following the same preference-to-score principle, we introduce PSDISTILL, a label-free self-distillation method for VLM aesthetic scoring. It converts VLM pairwise judgments into continuous pseudo-scores via an Elo reference pool, and trains the same VLM with ranking optimization with pair weighting to produce a single-pass aesthetic scorer. Trained on a single painting category, the distilled Qwen3-VL-8B improves mean SRCC from 0.504 to 0.709 across all three categories. Among the evaluated models, it achieves a higher mean SRCC point estimate than the specialized IAA baseline and most closed-source VLMs, using a single forward pass per image. Further experiments on APDDv2 assess the training method’s generalization beyond Chinese painting. We will release the training code and provide controlled academic access to the complete PPAINT annotations and score mappings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.