acceptodds
Under review as a conference paper at ICLR 2027

Read, Then Retrieve: Reader Feedback for Budgeted Multimodal Document Search

Abstract

Multimodal document retrieval typically applies an expensive reader to a fixed set of candidate pages, leaving relevant evidence outside that set inaccessible to reranking. We propose Reader Feedback, a two-batch acquisition method that uses intermediate reader judgments to select subsequent pages under a fixed reading budget. A frozen 7B reader first evaluates 16 visual candidates. A lightweight selector combines their relevance scores with semantic similarity, document membership, and uncovered query terms to choose 16 additional pages from visual and text retrieval pools. The same reader and a shared ranking head then order all 32 pages. Across three ViDoRe V3 development corpora, the fraction of queries with every annotated relevant page acquired increases from 55.34% to 62.43% over visual top-32 acquisition. On a separate HR corpus evaluated after method freezing, this coverage increases from 49.69% to 62.58%, while nDCG@10 rises from 64.97 to 67.75. Matched controls show a 5.97-point coverage gain over uniform feedback on HR, but the nDCG advantage over static hybrid retrieval is 0.13 points with an interval spanning zero. These results identify evidence acquisition as a useful role for reader feedback while distinguishing coverage gains from improvements in final ranking.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.