SQuAR: Synthetic Queries as Join-Aware Retrieval Anchors for Text-to-SQL
Abstract
Schema retrieval is critical in text-to-SQL systems over large database collections, where a system must identify both the target database and the tables needed to answer a natural-language question. Existing approaches either rely on costly query-time LLM reasoning or struggle to preserve multi-table structural relationships as the schema pool scales. To address this limitation, we present SQuAR, an open-pool schema retrieval approach that uses synthetic questions as join-aware retrieval anchors. Offline, SQuAR samples multi-table sub-schemas along foreign-key paths and indexes schema-grounded synthetic questions as retrieval anchors. At query time, it aggregates database-level and table-level evidence from retrieved anchors without query-time LLM calls. We further show that table-level recall alone is insufficient, as missing even one required table or selecting an incorrect database makes the retrieved schema insufficient to construct the target SQL query. Thus, we introduce Execution-Ready Coverage (ERC), which requires the correct database and all required tables, and Execution-Ready Jaccard (ERJ), which additionally penalizes over-retrieval. We evaluate SQuAR on four benchmarks and their unified pool of 37 databases and 619 tables. Across the four benchmarks, SQuAR improves Table Recall and ERC by up to 17.5 and 26.1 percentage points over the strongest evaluated baselines, while also maintaining a median retrieval latency of approximately 0.1 seconds. In the unified open pool, SQuAR improves DB R@1, Table Recall, and ERC by 15.5, 8.7, and 20.1 percentage points, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.