acceptodds
Under review as a conference paper at ICLR 2027

ROSA-IQA: Learning with Reasoning for Image Quality Assessment via On-Policy Self-Distillation

Abstract

Image quality assessment (IQA) takes two complementary forms: quality scoring quantifies perceptual quality, while quality description reveals the attributes and defects underlying human judgments. Although multimodal large language models (MLLMs) show promising reasoning capabilities, adapting them for quality description remains constrained by the fact that widely used IQA datasets typically provide only a scalar score per image. On-policy reinforcement learning directly learns from quality scores but provides only coarse sequence-level credit, whereas off-policy dense distillation from external MLLMs introduces a distribution shift between teacher demonstrations and student rollouts. We introduce a Reasoning-based On-policy Self-distillation Approach for IQA (ROSA-IQA). Under a teacher-student distillation paradigm, ROSA-IQA uses a frozen snapshot of the MLLM being optimized as the self-teacher, conditioned on curated privileged quality scores and segmentation-derived region guidance, to provide token-level supervision for student-sampled reasoning rollouts. We further use factual–counterfactual disagreement to identify tokens sensitive to quality conditions, thereby emphasizing quality-relevant descriptions in the student's reasoning. A scoring head then maps the combined visual and textual representations to continuous quality predictions. Extensive experiments show that ROSA-IQA outperforms leading no-reference IQA (NR-IQA) models in quality scoring while generating human-aligned quality descriptions. Importantly, we show that learning with reasoning improves cross-dataset generalization, suggesting that it acts as a regularizer that encourages the model to learn transferable quality features.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.