acceptodds
Under review as a conference paper at ICLR 2027

RStar-SQL: Self-Evolving Process Preference Learning for Complex-Reasoning Text-to-SQL

Abstract

Complex reasoning Text-to-SQL requires models to jointly resolve schema grounding, implicit constraints, domain knowledge, and SQL construction. While recent MCTS-based methods improve test-time exploration, their search feedback is typically used only for terminal candidate selection, providing limited reusable supervision for subsequent reasoning. We introduce RStar-SQL, a self-evolving process preference framework that converts SQL search outcomes into persistent training signals. RStar-SQL performs schema-grounded MCTS to generate diverse reasoning trajectories and evaluates them using execution consistency together with SQL-specific semantic signals, including schema grounding, execution behavior, join validity, and type operator compatibility. Verified high-quality trajectories refine the policy model, while preferred and dispreferred trajectory pairs train a SQL Process Preference Model. The learned SQL-PPM then guides action selection in subsequent MCTS rounds before terminal feedback is available, forming an iterative loop of search, verification, and policy preference refinement. Experiments on LogicCat and Archer show that RStar-SQL achieves execution accuracies of 35.41% and 47.82%, respectively, demonstrating strong effectiveness on complex reasoning tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.