acceptodds
Under review as a conference paper at ICLR 2027

Quality Signal Is Already There: Training-Free Cross-Domain IQA by Amplifying Latent Quality Direction in Frozen CLIP

Abstract

Cross-domain image quality assessment (CD-IQA) typically requires collecting target-domain images, aligning source and target distributions, and retraining a model whenever the deployment domain changes. Frozen contrastive vision–language models such as CLIP provide a training-free alternative, but their semantic-dominant representations are relatively insensitive to perceptual quality variations. By probing frozen CLIP across heterogeneous IQA datasets, we find that independently estimated quality directions exhibit non-random alignment and concentrate in a few common components, while the associated quality signal occupies only a small fraction of the overall feature energy. Perceptual quality is therefore latent rather than absent: transferable quality structure is already encoded in CLIP, but is underweighted in its default similarity space. This partial cross-domain commonality further suggests that a direction instantiated from an accessible synthetic source can serve as a transferable prior for strengthening quality-related components in unseen domains, without reproducing each target distribution. Based on this observation, we introduce a simple quality-directional feature amplification framework. A source-instantiated quality direction is constructed by contrasting high and low quality feature centroids, either using subjective ratings when available or, notably, without MOS by ranking synthetic distortions according to their affinity to pristine references. At inference, we amplify both the image-text feature using this source-derived quality prior before the standard cosine-similarity readout. The resulting method requires no target-domain data, parameter optimization, or additional network modules. Despite being training-free, it remains competitive with trained conventional cross-domain and VLM-based IQA methods, while consistently improving the frozen CLIP baseline across synthetic, held-out algorithmic, authentic, and AIGC quality benchmarks. Further analyses show that amplification strengthens distortion sensitivity and quality discrimination while largely retaining CLIP's semantic organization. The core code can be found in the Supplementary Materials.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.