Reasoning and Learning in Text-to-SQL: From Template Imitation to Strategy Recombination
Abstract
Large language models increasingly use chain-of-thought (CoT) reasoning to translate natural-language questions into SQL queries. Existing methods provide diverse reasoning templates, but how models use these templates after supervised fine-tuning and reinforcement learning remains unclear. In particular, can learning from multiple templates lead to useful reasoning beyond the imitation of a single template? In this work, we empirically study this transition on BIRD using two 7B models. We construct matched training data with balanced coverage of four templates and compare supervised fine-tuning, rejection fine-tuning, and subsequent GRPO. By analyzing freely generated responses, intervening on requested templates, and comparing reasoning functions before and after RL, we find that balanced training does not produce balanced template use. Instead, RL changes both the dominant forms of expression and the organization of reasoning functions within complete responses. Among same-question greedy responses with identifiable reasoning functions, RL changes their combination or order in 78.74% and 73.11% of cases for the two SFT models. Reasoning organizations associated with successful solutions become more frequent after RL across all four supervised training paths, including when these organizations are identified using pre-RL responses alone. These findings show how multi-template learning and RL can support a shift from template imitation toward task-useful CoT organization, providing an empirical basis for studying strategy recombination in text-to-SQL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.