Exact, Batch-Invariant and Fast Test-Time Query Optimization for Late-Interaction Retrieval
Abstract
Test-time query optimization (TTO) refines a query embedding by gradient steps against a frozen candidate pool and lets compact multi-vector retrievers rival much larger ones. With late interaction, every step rescans every vector of every candidate, and how these scans are executed changes the result: TF32 tensor cores change maxima and selected vectors, and library matrix products change their bits with the batch a request is served in. We present ExactTTO, which makes TTO exact, that is, bitwise identical to a fixed FP32 reference, batch-invariant and fast by letting approximation decide only which products are computed, never their values, and fixing the order of all other arithmetic. Certified working sets skip rescans that provably cannot change the result. Tensor cores only filter candidates, whose FP32 maxima and selected vectors are recomputed exactly. A fused loop applies each step's gradient and Adam update in one kernel with a fixed operation order, so a query's trajectory does not depend on the requests it is batched with. Under the official protocol of Guided Query Refinement (GQR) on fourteen ViDoRe splits and four BEIR datasets, we recover GQR's gains. In batches of 64 requests, ExactTTO processes 3.05× as many queries per second as GQR's FP32 scorer in our fused loop and 3.95× as many as FLASH-MaxSim in FP32, and gives every query the bits it gets when served alone, which GQR's scorer does for 80.4% of the queries; a single ColNomic refinement runs 21× faster than GQR's released eager loop, mostly by removing framework overhead. ExactTTO also accelerates a second objective, distilling a cross-encoder into ColBERTv2 queries, which raises nDCG@10 on BEIR from 37.1 to 41.5.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.