Attribute-Calibrated and Quality-Aware Vision-Language Learning for Image Quality Assessment
Abstract
Existing image quality assessment (IQA) methods primarily rely on score regression paradigms while neglecting semantic analysis, making score-only supervision unreliable across diverse scenarios. Additionally, the relationship between quality-aware attributes and cross-scenario quality fluctuations remains underexplored. To address these issues, an image quality assessment method based on attribute-calibrated and quality-aware vision-language learning (ACQA) is proposed. Firstly, a dual-branch vision-language model (VLM) is employed to extract content and attribute features for subsequent quality analysis. Secondly, an attribute calibration learning (ACAL) module is proposed to calibrate the VLM in attribute analysis by jointly incorporating distortion ranking, stability, and disentanglement, thereby enabling adaptive attribute information across scenarios. Thirdly, a quality-aware analysis learning (QAL) module is proposed to integrate attribute contributions with content alignment to optimize the VLM. Finally, the optimized VLM provides reliable content and attribute information for quality prediction. Experimental results demonstrate that the proposed ACAL and QAL can effectively refine content and attribute information to enhance the reliability of quality prediction. Moreover, ACQA outperforms existing IQA methods in both accuracy and generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.