DeltaBind: Relation-Bound Retrieval for Visually Rendered Tables
Abstract
Visual table retrieval can fail when candidates contain the same words but associate values with different rows or columns. Page-wide late interaction can combine query-token matches from incompatible relations, a failure we call relation collision. We introduce DeltaBind, a training-free framework that makes a complete table relation the unit of query matching. Using predicted table structure, DeltaBind assembles source-image crops of a value, its row label, and its column header into a relation tile. A frozen retriever scores each tile with late interaction and ranks a table by its highest-scoring tile, keeping query-token matches within a shared relation. For aligned table versions, shared content can obscure the relations that distinguish candidates. A query-independent change index therefore focuses comparison on changed relations and enables sparse caching. DeltaBind improves ColQwen2's strict pair accuracy from 51.6% to 79.7% on TabFact-HumanBind. Relation binding alone also improves source-local retrieval on mined pools of unedited FinTabNet and PubTables-1M tables. OCR-BM25 controls and component ablations support the value of jointly preserving row, column, and value context.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.