Human-Level Text-to-SQL via RL with Verified Data and Task-Specific Rewards
Abstract
Despite increasingly complex multi-stage inference pipelines, existing text-to-SQL systems are still limited by the SQL reasoning capability of their underlying large language model (LLM). These systems still underperform human experts by over 10% on benchmarks. In this paper, we show that reinforcement learning with verified training data and task-specific rewards enables a single fine-tuned LLM to achieve human-level performance. We introduce BIRD-Platinum, a dataset of 2.5k instances sampled from BIRD Train and curated through multiple rounds of expert verification. We corrected annotation errors in 61% of the instances. We then diagnose two limitations of standard execution-based outcome rewards. First, they can reward semantically incorrect queries that coincidentally return the correct result. Second, they provide no direct incentive to use provided external knowledge. To address these limitations, we propose ReViSQL, which combines execution-based grading with bounded SQL equivalence verification and process rewards that encourage external-knowledge use. We show that fine-tuning Kimi-K2.6 on BIRD-Platinum with REVISQL matches the 92.96% human performance on Arcwise-Plat, the expert-verified variants of BIRD. Our model outperforms open-source pipelines by 9–22% and surpasses GPT 5.6 Sol Ultra and Claude Fable 5 by 3–8% at 12–15% of their inference cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.