Internalizing Execution for Text-to-SQL with Predictive Relational Verification
Abstract
Execution feedback helps Text-to-SQL systems identify and revise incorrect query choices, but obtaining it typically requires interaction with a database engine. We investigate whether models can provide this feedback by predicting intermediate query results during SQL construction. On the same BIRD questions, we compare schema-only SQL generation with prediction of the corresponding gold SQL's results on MiniDBs constructed using that SQL. Models can accurately predict these results yet struggle to generate correct queries for the same questions, suggesting a gap between execution capability and use. We introduce SimSquirrel, a family of 4B, 8B, and 14B models trained to use execution predictions during SQL construction. Within a single continuous response, each model generates local SQL, predicts its result relation, and selects Continue or Fix to guide query construction without external execution feedback. To supply data for these predictions, we construct diagnostic MiniDBs without reference SQL, selecting linked rows within the context budget. We train this behavior by first developing execution prediction through supervised fine-tuning and reinforcement learning, then optimizing complete Text-to-SQL trajectories. During trajectory training, identical rewards on a single database yield zero group-relative advantages. Our counterexample-guided dynamic sampling therefore attempts to distinguish candidates from the reference query on shared synthesized databases, rescores each group, and retains groups with differing rewards for policy optimization. On BIRD dev, SimSquirrel-14B achieves 73.86% execution accuracy, exceeding Qwen3-14B by 23.99 percentage points and the strongest baseline in our comparison by 3.06 points. Training-stage ablations show that, at 8B, the complete method exceeds Text-to-SQL reinforcement learning alone by 12.79 points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.