acceptodds
Under review as a conference paper at ICLR 2027

Quality Has a Direction: Learning a Shared Quality Axis from Reasoning for Image Quality Assessment

Abstract

Reinforcement-learning-based vision–language models generate quality reasoning that provides effective supervision for generalized image quality assessment, yet the resulting text embeddings entangle perceptual quality with scene content and linguistic realization. We investigate whether transferable quality can be localized within this entangled space. We find that MOS-conditioned variation in reasoning text forms a semantically coherent quality axis despite contextual and lexical ambiguity. Analyses of contrastive training dynamics and controlled domain shifts further reveal that the quality axis is shared by language and vision across datasets. Building on this finding, we propose QAxis, which introduces the shared quality axis into contrastive learning. Symmetric RoPE makes cross-modal similarity jointly depend on semantic correspondence in the complete representations and agreement along their quality coordinates, while axis-wise ordinal learning establishes a global perceptual order. Using only a 0.1B-parameter image encoder, direct axis projection achieves state-of-the-art average PLCC/SRCC of 0.811/0.796 across seven benchmarks; incorporating the complete image representation further improves the averages to 0.816/0.800.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.