acceptodds
Under review as a conference paper at ICLR 2027

A Text-Parametrized Metric for Image Similarity

Abstract

Although assessing visual similarity comes naturally to humans, measuring such perceptions remains a challenge. State-of-the-art metrics typically only consider one approach or format for measuring similarity, and do not generalize well to new scenarios. While recent methods begin to compare images in ways beyond overall visual appearance, there are many kind of similarities without an established metric. For example, if users want to compare images with respect to art style, noise pattern, or even the number of cats, would they need to construct a dataset and learn a metric from scratch? Instead, our goal is to learn a single model where the similarity context, in addition to the images, is also a model input. To this end, we learn a text-parametrized perceptual metric that predicts conditional similarity specified in the text prompt. We construct a two alternative forced choice (2AFC) dataset to finetune a perceptual embedding where image representations are influenced by the comparison described in the text. This allows us to compare fine-grained and complex visual attributes with one model. We study our model's capability to measure visual similarity in diverse contexts, demonstrate applications on conditional image retrieval, and compare to existing similarity metrics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.