Toward Unified Multimodal Feature Quality Assessment via Cross-Modal Feature Alignment
Abstract
For decades, perceptual quality assessment for image, text, and audio relied on three separate toolchains with incompatible metrics and almost no overlap. Yet the common ground is deeper than it appears. If we can project features from all three modalities into the same space, quality assessment becomes a single regression task. To test this, we built a synthetic distortion benchmark covering all three modalities with consistent labels, all traced back to shared visual references. We designed a lightweight pipeline that first identifies the modality and then applies a dedicated regressor, using a simple mechanism to bridge the feature spaces. The result is quality assessment across modalities without heavy custom engineering. Our experiments find a consistent difficulty ordering, clear differences in how models regularize, and a hard limit on how well standard audio features transfer to unseen distortions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.