acceptodds
Under review as a conference paper at ICLR 2027

ReprForge: Efficient Index Rebuilding for Visual Document Retriever Upgrades

Abstract

Upgrading a visual document retriever usually requires re-encoding the entire collection. However, this process repeats costly visual computation, which may remain valid after the upgrade. Reusing such computation requires determining which visual results the target model needs and whether they remain valid. In this paper, we present ReprForge, which answers these questions by tracing how visual outputs are produced and consumed. From one forward pass, it identifies the complete visual state required by later stages and its computational dependencies, then retains this state during index construction. At upgrade time, it checks the dependencies as loaded and the processed page inputs. When both remain unchanged, the target model resumes from the retained state with newly generated text embeddings and positions; otherwise, it performs full encoding. This links reuse to the actual computation rather than a manually chosen cache boundary. On 76 released upgrade pairs, reuse decisions agree with independent bitwise checks on tested pages. On 20,946 pages under a common processing configuration and a fixed input contract, four real upgrades achieve 4.35–4.48× rebuild speedups and reduce total build and rebuild time, including the initial build and state retention, by 61%, with bitwise-identical indexes under the tested settings. These results show that exact index updates do not require repeating the full encoding process: intermediate visual computation can remain reusable even when the final document representations change.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.