Diagnosis-Score Decoupling for Cross-Dataset Image Quality Assessment
Abstract
Vision-language models (VLMs) can describe image degradations, yet their quality predictions remain sensitive to dataset-specific scoring conventions and distortion distributions. It remains unclear whether chain-of-thought (CoT) diagnostic descriptions provide transferable information beyond numerical quality prediction. This paper introduces Diagnosis-Score Decoupling (DSD), a two-branch framework for CoT-based image quality assessmnet (CoT-IQA) that explicitly separates diagnostic evidence from quality scoring. One branch predicts a scalar quality prior, while the other generates structured CoT diagnostics describling degradation type, severity, and spatial extent. Rather than using the diagnostic branch’s final score, DSD parses its descriptions into a compact, score-free representation. A lightweight calibrator then combines this representation with the scalar prior using only source-domain quality labels and is transferred unchanged to target datasets without target-domain labels, tuning, or recalibration. Across all six evaluated transfers between controlled-distortion datasets, DSD consistently improves SRCC over the scalar-prior baseline (MOS only) on the primary backbone, with a similar overall trend on a second VLM backbone. Ablation studies indicate that the gains driven primarily by image-aligned diagnostic cues, particularly degradation severity. On KonIQ-10k, however, these gains do not persist, suggesting that the benefit of source-calibrated diagnostic evidence depends on the target distortion regime.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.