SAGA: Single-Sequence Coupling of Description and Scoring via Scene-Attribute Graphs based Aesthetic Learning
Abstract
Image quality assessment has evolved from a simple numerical regression to integrate score prediction with explanatory text generation. Existing methods usually treat these two objectives as completely independent tasks and train them in a decoupled manner, which may break the inherent causal relationship between the two tasks. Yet, directly modeling both tasks within a single generation sequence often leads to optimization conflicts and training instability due to the heterogeneity of target formats and imbalanced supervision signals. In this paper, we propose a novel Scene-Attribute Graphs based Aesthetic learning (SAGA) framework. Dual expert branches are fused within SAGA, based on their task-specific aesthetic experience latent spaces, to enhance the fusion model's capacity of capturing intricate aesthetic preferences. Thus, SAGA can effectively unify description generation and quality scoring within a single sequence. Extensive experiments show that SAGA significantly outperforms the state-of-the-art methods on both description generation and quality scoring, while exhibiting more balanced multitask performance. Furthermore, compared with models trained separately, SAGA maintains strong reasoning capability, while achieving notably lower inference latency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.