Generative Late-Interaction Embeddings For Visual Document Retrieval
Abstract
Late-interaction retrieval is the state-of-the-art for visual document search, but storage costly, where compression techniques retain a subset or local average of the vectors per page. Under aggressive memory budgets, these methods degrade sharply, while alternatives require retraining the encoder. Investigating this degradation across three encoders revealed two geometric properties: vectors lie on the unit sphere and concentrate within a 5-6 dimensional manifold. Consequently, a) standard -means centroids fall inside the sphere, underestimating MaxSim scores and normalizing them to the surface fixes this ( nDCG@5). b) the low-dimensional manifold allows full vector regeneration from a few. To this end, we introduce Generative Late-Interaction Embeddings (GLIE): vectors per page learned from the normalized centroids to serve as both a lightweight index and a basis for regenerating a page's full embedding set. At query time, search runs exclusively on these vectors, and a decoder expands only the top candidates back to all vectors for exact re-scoring. At four vectors per page on ViDoRe v1, GLIE retains nearly 80% of the uncompressed system's nDCG@5, against 70% for the best prior post-hoc method. These results use a 415K-parameter network fitted in within three GPU-minutes on just a few thousands training pages. At a matched training budget, fine-tuning the encoder does not reach even the training-free stage of GLIE, and the full system beats it at every budget. These patterns are true for a second encoder and ViDoRe v2. By reconstructing evidence on demand rather than sampling it, GLIE opens a new axis for storage-efficient retrieval, with the decoder as its main design surface.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.