Is One Direction Enough? Visual Embeddings Rankability for Image Quality Assessment
Abstract
Image quality assessment (IQA) models, from classic deep neural networks to multimodal large language models (MLLMs), are evaluated by how well their predicted scores correlate with human mean opinion scores (MOS). However, this evaluation does not reveal that the hidden representations are rankable by quality scores. In this work, we study representations derived from hidden layers of IQA models and propose to estimate a single linear direction, the rank-axis, and measure whether projections onto it order images by quality scores (rankability). Additionally, we propose the Rank-Axis Isometry Score (RAIS) to quantify how well their spacing reflects the size of quality differences. To test whether the embedding space itself is ordered by quality scores, we introduce three quality-embedding alignment measures computed on distances in the native feature space. Across nine IQA models, three layer depths and eight benchmarks, one rank-axis reaches a rankability of at least on six datasets, and at least one of the rank-axes across hidden representations beats the model's own fine-tuned head in 52 of 70 encoder-dataset pairs. Spacing along the axis follows the order (Spearman between RAIS and rankability), whereas the embeddings in the original latent space are only weakly ordered by quality (mean Anchor-QEA ). Rank-axes estimated on different datasets are nearly orthogonal, yet keep 80% of their rankability on unseen datasets; their agreement under the data covariance predicts this transfer (Spearman ). Our findings suggest that a single linear direction captures much of the perceptual quality ordering in visual embeddings, even when it is not reflected in their dominant geometry.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.