Database Context Compression for Text-to-SQL on Real-World Large Databases
Abstract
Recent progress on Text-to-SQL has been driven by stronger language models and richer prompting strategies, yet performance on real enterprise benchmarks such as Spider 2.0 and BIRD remains far below that on classical academic datasets. We argue that a central bottleneck on such benchmarks is the way the database is presented to the model. Real databases contain wide tables with repeated audit columns, large families of homogeneous partitioned tables, opaque machine-generated identifiers whose meaning lives only in column descriptions, and long data dictionaries in which only a small, query-dependent fraction is actually relevant. Existing query-aware approaches—schema linking, broadly construed to include retrieval-style schema subsetting—attempt to filter this raw context, but often operate on a representation that is structurally redundant, semantically verbose and documentation-heavy. We re-frame the problem as one of database context compression: a database-side rewrite that increases the information density of schema structure and semantic descriptors, complemented by query-aware purification of external documents. We introduce the SGCF (Support–Gain Component Factorization) principle as a common accounting abstraction for repeated column groups, isomorphic table templates and shared semantic-tag components; question-relevant evidence purification addresses the distinct, query-dependent document layer. Building on this abstraction, we present DBCC, a two-phase database-side middleware that performs query-agnostic structural and semantic re-encoding offline, and lightweight query-aware evidence purification online. DBCC is model-agnostic and pipeline-agnostic, and can be inserted before schema linking or generation in existing Text-to-SQL systems. On Spider 2.0-Snow and BIRD, DBCC reduces input tokens by up to 75x (on the most challenging Large bucket of Spider 2.0-Snow, from 2.6 M to 34.7K tokens). On that bucket, where the raw input is over budget and produces no completed predictions, the compressed input reaches 56.5% strict recall under DeepSeek-V3.2 and 63.1% under Claude-Opus-4.7. DBCC also yields a 1.8-1.9 percentage-point end-to-end EX improvement when stacked on top of each of three recent Text-to-SQL systems. Code and evaluation artifacts are available in the anonymous repository at https://anonymous.4open.science/r/SchemaCompression/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.