acceptodds
Under review as a conference paper at ICLR 2027

Latent Aesthetic Deliberation: Learning to Refine Aesthetic Judgments

Abstract

Image Aesthetic Assessment (IAA) requires weighing visual cues whose aesthetic contributions depend on image context. Although image–rating pairs are widely available, they specify judgment outcomes rather than how evidence should be integrated or assessments revised. Direct scoring with multimodal large language models (MLLMs) does not explicitly model judgment revision, whereas supervised verbalized reasoning requires additional reasoning-text targets and incurs token-by-token decoding overhead. Motivated by this, we introduce Latent Aesthetic Deliberation (LAD), a recurrent framework for learning judgment refinement from ratings without annotated reasoning trajectories. After a single multimodal encoding pass, an evolving judgment state repeatedly retrieves and integrates evidence from persistent image-token memory. Regression against human ratings trains the recurrent scorer, while internal self-distillation aligns intermediate score readouts with the final prediction without prescribing their latent representations. At inference, LAD fuses the recurrent score with a parallel language-aligned rating estimate without decoding an explanation. Extensive experiments demonstrate state-of-the-art performance on the evaluated benchmarks for holistic and aspect-specific aesthetic assessment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.