acceptodds
Under review as a conference paper at ICLR 2027

Toward Unified Multimodal Feature Quality Assessment via Cross-Modal Feature Alignment

Abstract

For decades, perceptual quality assessment for image, text, and audio relied on three separate toolchains with incompatible metrics and almost no overlap. Yet the common ground is deeper than it appears. If we can project features from all three modalities into the same space, quality assessment becomes a single regression task. To test this, we built a synthetic distortion benchmark covering all three modalities with consistent labels, all traced back to shared visual references. We designed a lightweight pipeline that first identifies the modality and then applies a dedicated regressor, using a simple mechanism to bridge the feature spaces. The result is quality assessment across modalities without heavy custom engineering. Our experiments find a consistent difficulty ordering, clear differences in how models regularize, and a hard limit on how well standard audio features transfer to unseen distortions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.