Query-Conditioned Residual Retrieval for Multimodal RAG
Abstract
Multimodal retrieval-augmented generation (RAG) extends external knowledge beyond text to heterogeneous sources such as images, documents, and videos. A practical strategy represents such content using query-independent textual descriptions, whereas recent multimodal retrievers directly index the original multimodal content. These alternatives expose a basic representation problem: a textual description can be semantically correct yet omit a visual or temporal detail required by a particular query. We study this problem through query-dependent description sufficiency and introduce Query-Conditioned Residual Retrieval (QCR), a training-free retrieval approach. For each multimodal item, QCR uses its textual descriptions to construct an item-specific description subspace in a shared multimodal representation space and retains the orthogonal component as a description-conditioned residual. At inference time, QCRindependently retrieves from the textual description and residual representations. The first route captures relevance directly expressed by the descriptions, while the second exposes representation directions not contained in their span. Candidates from both routes are jointly ranked, and the selected original items are provided to the answer model. Across heterogeneous text, table, image, and video knowledge sources, QCR consistently improves over description-only, direct multimodal, and description–multimodal fusion retrieval. Extensive analysis shows that QCR advantage grows as textual descriptions become less sufficient for the downstream query, together with increased recovery of supporting evidence missed by description retrieval.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.