From Pipeline-Based Generation to Interactive Reasoning Trajectory Learning for Text-to-SQL
Abstract
In the intersection of natural language processing (NLP) and structured databases, translating natural language queries into SQL remains a formidable challenge. While current methods, which combine multi-stage pipelines with LLMs, have made substantial progress in generating executable SQL, they still encounter issues such as error propagation and high generation latency. To address these challenges, we propose STIR-SQL, a Text-to-SQL agent that performs single-trajectory interleaved reasoning and is trained via pure reinforcement learning with multi-dimension reward. Specifically, STIR-SQL replaces the multi-stage pipeline with a single, tool-integrated interleaved reasoning trajectory before the final answer, directly reducing generation latency. Within this trajectory, STIR-SQL interacts with specialized tools and performs self-correction, which mitigates error propagation and enhances the accuracy of SQL generation. Experimental results show that our approach achieves an execution accuracy of 69.7% on the BIRD benchmark, while requiring only 3B activated parameters during inference in a 30B-scale model. Furthermore, the reduction in activated parameters and inference latency demonstrates its strong potential for real-world deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.