PACE: Mitigating Autoregressive Scoring Biases for Robust Image-Text Retrieval
Abstract
Multimodal large language models (MLLMs) offer rich cross-modal reasoning capabilities for image-text retrieval, but their high inference cost and biased gen- erative likelihood limit their scalability. Global dual-encoder scores also struggle to distinguish highly similar negative samples. We propose PACE (Progressive Adaptive Cross-modal Enhancement), a training-free coarse-to-fine framework that selectively introduces MLLM-based evidence into efficient retrieval. PACE first employs SigLIP2 to construct a compact candidate pool and uses a margin- based guard to route only ambiguous cases to MLLM refinement. For large-scale reranking, we calibrate image-conditioned likelihood with a text-only prior us- ing pointwise mutual information (PMI), reducing the influence of language pri- ors. For fine-grained matching, we combine Position-Discounted Autoregressive Scoring (PDAS), which discounts later token contributions, with a Contrastive En- tailment Margin (CEM) based on visual Yes/No verification. The resulting frame- work separates efficient candidate retrieval from fine-grained multimodal verifica- tion, allowing MLLM computation to be allocated where it is most informative. Experiments across image-text retrieval and compositional benchmarks demon- strate improved fine-grained matching while substantially reducing unnecessary MLLM inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.