acceptodds
Under review as a conference paper at ICLR 2027

Faithful Textual Explanations for Interpretable Image Aesthetic Assessment

Abstract

In image aesthetic assessment (IAA), explanations typically rely on evaluative terms like harmonious or nostalgic. However, these terms describe the overall impression rather than specific regions of an image, and they lack ground-truth annotations for evaluation. Existing concept bottleneck models (CBMs) cannot be easily applied here because they rely on localized, verifiable visual concepts. Furthermore, selecting concepts based solely on their correlation with aesthetic ratings often fills explanations with paraphrases of "good image" or with descriptions of basic image content rather than aesthetic quality. To address these limitations, we introduce FaTeX, an interpretable IAA model that generates faithful, grounded, and image-specific textual explanations without requiring manual concept annotations. FaTeX calculates aesthetic scores by summing contributions from named concept pairs generated by a large language model (LLM) and scored using CLIP. It controls for image content, filters out redundant paraphrases, and uses concepts that apply both across datasets and within specific content categories. A gating step is designed to give no weight to level-specific concepts whose subject does not appear in the image, so that they drop out of the explanation. Evaluated across six single- and multi-genre datasets, FaTeX trails the best non-interpretable models by .035 SRCC on average (at most .058) and outperforms existing attribute-based approaches. In human and LLM-based evaluations measuring explanation quality and image identification, FaTeX produces better and more specific explanations than existing concept-based methods, and is on par only where a baseline draws on expert style annotations. Code and materials are available at https://anonymous.4open.science/r/proj-FaTeX-E4FB.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.