Don't Generate Just to Discard: Prompt-Side Aesthetic Forecasting for Text-to-Image
Abstract
Achieving satisfactory text-to-image (T2I) results often requires costly trial and error: users repeatedly revise prompts, generate images, and evaluate them manually or with quality assessment models, wasting computation and waiting time on unsatisfactory outputs. This leaves a practical but underexplored question: can we assess a prompt's quality potential before generating any images? Existing prompt-side studies find meaningful correlations between prompt embeddings and image quality averaged across repeated generations, suggesting that expected generation quality can be partially forecast from text alone. Yet limitations remain at three levels. At the quantification level, they collapse stochastic outcomes into an average score, obscuring variation across images generated from the same prompt. At the modeling level, coarse prompt-to-score mappings do not capture contextual interactions or address the cost of obtaining reliable multi-sample supervision. At the interpretability level, average attribution makes it difficult to identify which prompt components improve or impair image quality. Motivated by these limitations, we formulate Prompt-Side Aesthetics Forecasting (PSAF), a pre-generation task that estimates a given generator's expected aesthetic response from prompt text alone, and instantiate it with AesPrompt. AesPrompt uses learnable Aesthetic Response Queries to distill distributional aesthetic evidence from multiple generated samples into prompt-side representations, together with a Chunk-Aware Prompt Adapter that captures contextual interactions in long prompts. We further contribute PSAF-540K, a dataset of 540K samples designed for the PSAF task. During dataset construction, adaptive Bayesian response sampling directs the limited generation budget toward prompts expected to provide the most informative supervision. With only 95.7M trainable parameters, approximately 0.3% of a 30B-scale generator, AesPrompt achieves a Spearman rank correlation (SRCC) of 0.951 and a pairwise ranking accuracy (PairAcc) of 0.944 without sampling images at inference time. In typical prompt exploration workflows, filtering with AesPrompt can reduce the number of rendered candidates and GPU time by 60%–85% while retaining high-quality prompts for final generation. The code, datasets, and demo are included in the supplementary materials.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.