acceptodds
Under review as a conference paper at ICLR 2027

MMProvence: Turning a Multimodal Reranker into a Zero-Cost Context Pruner

Abstract

Multimodal RAG pipelines pass the retrieved pages to a vision-language model (VLM) as images together with their text, and the page images make up most of the model's input, yet existing context pruners operate on text only. We introduce MMProvence, a zero-cost context pruner for multimodal RAG: a VLM reranker that, in a single call per query-page pair, generates a relevance score, a bounding box on the page image, and the page sentences that answer the query. The score reranks the candidates, and the box and the sentences prune the image and the text of the pages that survive reranking, so pruning requires no additional model call. We further show that what the box is trained to cover decides the cost of cropping: on top of text pruning, crops to boxes that include the context that makes the evidence readable, such as table headers or chart legends, lower the LLM-as-a-judge score by 2.7 points, against 13.3 points for boxes around the smallest answer-bearing region. On the eight domains of ViDoRe V3, MMProvence-2B reduces the generator's input by 75%, from 21.6 to 5.3 tokens, at an LLM-as-a-judge score of 0.803 against 0.827 with the full context, while ranking on par with the reranker it replaces (nDCG@10 of 0.601 vs. 0.604). At 4B and 9B parameters, MMProvence ranks better than this reranker (0.627 and 0.636), and MMProvence-9B keeps a score of 0.822 while removing 76% of the input. We will release our code, checkpoints and training data, together with VERGE, our framework for end-to-end multimodal RAG evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.