Content-Adaptive Aesthetic Alignment for Text-to-Image Generation
Abstract
Aesthetic evaluation for text-to-image (T2I) generation must distinguish an inap- plicable criterion from a poorly realized one. To assess diverse visual content, we define a taxonomy of six aesthetic criteria, *i.e.*, realism, lighting, anatomy, creativ- ity, color harmony, and typography, and construct an aesthetic preference corpus of approximately 240K distinct images and 840K pairwise comparisons from hu- man and GPT-assisted criterion-specific rankings. Based on this corpus, we train **A**pplicability-Aware **W**eighted **A**esthetic **RE**ward (AWARE), which jointly pre- dicts quality and applicability for all six dimensions in one model. We then use AWARE to design an applicability-aware reinforcement-learning algorithm for text-to-image post-training, dynamically selecting and weighting only the crite- ria relevant to each sample. To evaluate both aesthetic judgment and post-trained generation, we introduce AWARE-RBench for criterion-conditioned preference and applicability evaluation, and AWARE-GenBench for role-conditional T2I aes- thetic quality. AWARE improves over the baseline by 60.92% on the aesthetic- model evaluation. Across Z-Image and SenseNova-U1.5, post-trained T2I models achieve average relative gains of 19.01%, 21.72%, and 2.39% on Photography, Artwork, and Infographic, respectively. On general-purpose benchmarks such as Qwen-Image-Bench, the aligned models remain competitive with existing open- source systems. Side-by-side visual comparisons further confirm improvements in fidelity, composition, natural lighting, and text rendering.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.