MINER: Mining Multi-Modal Internal Representation for Efficient Retrieval
Abstract
Visual document retrieval has become essential for accessing information in visually rich documents. Existing approaches fall into two camps. Late-interaction retrievers achieve strong quality through fine-grained token-level matching but store hundreds of vectors per page, incurring large index footprints and high serving costs. By contrast, dense single-vector retrievers retain storage and latency advantages but consistently lag in quality because they compress all information into a single final-layer embedding. In this work, we first conduct a layerwise diagnostic on single-vector retrievers, revealing that retrieval-relevant signal resides in internal representations. Motivated by these findings, we propose **MINER** (**M**ining Multimodal **I**nternal Represe**N**tation for **E**fficient **R**etrieval), a lightweight plug-in module that probes and fuses internal signals across transformer layers into a single compact embedding without modifying the backbone or sacrificing single-vector efficiency. The first *Retrieval-Aligned Layer Probing* stage attaches a lightweight probe at each layer, surfacing which dimensions carry retrieval-relevant information. The subsequent *Adaptive Sparse Multi-Layer Fusion* stage applies performance-adaptive neuron-level masking to the selected layers and fuses the surviving signals into the final dense vector. Across ViDoRe V/V/V, MINER outperforms existing dense single-vector retrievers on the majority of benchmarks, with up to 4.5% nDCG@ improvement over its corresponding backbone. In settings where models support both late-interaction and dense retrieval, MINER substantially narrows the nDCG@ performance gap while retaining the storage and serving efficiency of dense retrieval. Our source code is available at [this anonymous repository](https://anonymous.4open.science/r/MINER-8F41).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.