Advancing the Text-to-SQL Frontier through Graph-Language Verification
Abstract
Translating natural language queries into Structured Query Language (SQL) in real world databases requires reconciling ambiguous user intent with complex relational schemas. While modern LLMs can generate candidate SQL queries with ease, identifying whether a query is functionally correct remains a fundamental challenge. Existing selection heuristics (e.g. execution-based self-consistency and LLM-as-a-judge evaluators) fail to handle the inherent syntax-semantics asymmetry of SQL. In other words, SQL has many ways to express equivalent queries, but small (single-token) variations can dramatically change a query's result. In this work, we formulate SQL candidate verification as a cross-modal alignment problem between unstructured textual context (user intent, schema definitions, and domain evidence) and structured relational computation graphs. We introduce Warbler, a unified framework coupling execution-guided candidate generation with a learned graph-language verifier which algorithmically compiles SQL into directed acyclic graphs (DAGs), normalizes syntactic variance, and encodes relational inductive biases via a graph-language architecture. On the highly competitive BIRD benchmark, Warbler establishes a new state of the art on the official single model track with 80.04% execution accuracy on the held-out test set, outperforming prior leading method by 2.90 points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.