acceptodds
Under review as a conference paper at ICLR 2027

QMORE: Robust Quantum Image–Text Retrieval via Encoder Sharing and Diversity-Aware Rank Fusion

Abstract

We present QMORE, a hybrid quantum–classical framework for cross-modal image–text retrieval. QMORE learns a shared image–text geometry by reusing a single variational quantum encoder across both modalities, paired with lightweight modality adapters that preserve modality-specific flexibility while keeping quantum parameters shared. To reduce run-to-run sensitivity, QMORE ensembles independently trained models, weights members by retrieval strength and diversity, and fuses predictions via weighted rank voting. Across three benchmarks, QMORE improves Recall@1 over the strongest parameter-matched baseline by , , and points on SVO Probes, Captioned CIFAR-10, and Flickr30k—a – relative improvement on every dataset—while also achieving the best mean and median rank throughout. The advantage persists against classical encoders with more encoder-core parameters, and under a calibrated noise model of a near-term IBM device QMORE retains of its Recall@1 while the quantum baseline collapses to near-random.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.