acceptodds
Under review as a conference paper at ICLR 2027

GESSO: An Aesthetic Perceptual Loss for Pixel Space Text-to-Image Pretraining

Abstract

Aesthetic quality is conventionally added through data filtering, curation, preference finetuning, or reward optimization. Here, we ask whether a perceptual loss can supply aesthetic signal during pretraining. We study two design choices not previously examined in text-to-image pretraining. First, a system-level study of ten ViT-B encoders shows that language-supervised encoders perform best, while the self-supervised default, DINOv2, is last or tied for last on aesthetics, prompt alignment and fidelity. Swapping in SigLIP gives a stronger perceptual recipe. Second, how an aesthetic scorer enters the loss matters: with the scorer fixed, distilling it into a reference-anchored, per-patch distance improves held-out aesthetic metrics, whereas maximizing its score does not. We introduce GESSO, an aesthetic perceptual loss for pretraining: a lightweight projector over frozen SigLIP features maps a reference image and the model output to a distance that tracks the difference in their aesthetic scores, so aesthetic signal reaches every pretraining image rather than only a small curated set afterwards, at 0.8% extra step time. Added to the SigLIP recipe, GESSO improves every human-preference metric and is preferred by every VLM judge across three model scales. At B/32 the full recipe improves HPSv2.1 by 10.2% and reduces FID by 12.4% relative to the PixelGen recipe; GESSO alone contributes 4.3% and 6.0%. Prompt alignment is similar, and the gains persist after supervised finetuning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.