acceptodds
Under review as a conference paper at ICLR 2027

When Tables Need Context: Query-Conditioned Table Extraction

Abstract

Most document table extraction systems follow an extract-all paradigm: they reconstruct every table in a document before downstream retrieval or filtering. However, real-world document interaction is often query-driven, where users seek only the tables relevant to a specific information need. We introduce Query-Conditioned Table Extraction (QCTE), a new task that identifies tables satisfying a natural-language condition and reconstructs only the matched tables as structured HTML. We argue that query-conditioned extraction can benefit from document-level context beyond isolated table crops, since captions and surrounding paragraphs provide useful evidence for reliable condition judgment. To this end, we propose PCRC, a framework that combines table-centered context retrieval, progressive evidence incorporation, and reusable KV-cache inference. PCRC progressively introduces contextual evidence only when necessary while sharing visual-textual representations across table selection and reconstruction, substantially reducing redundant multimodal computation. To support this setting, we construct three condition-grounded benchmarks spanning financial, scientific, and visually degraded documents. Across all three datasets, PCRC improves judgment F1 from about 0.73 to 0.90 and composite F1-TEDS by 13-16 points. It requires only 24-29% of the prefill computation of the same inference schedule without cache reuse. These results show that selective document context improves query-conditioned table extraction and can be used efficiently with cache reuse.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.