SQL-Structure-Aware On-Policy Distillation for Text-to-SQL
Abstract
On-policy distillation (OPD) provides token-level teacher feedback for training Text-to-SQL models. However, uniform weighting allows lengthy reasoning text to dominate the distillation objective and does not account for how tokens jointly express SQL operations, columns, and values. We propose a SQL-structure-aware weighting method for OPD that controls the total weight assigned to final SQL and reasoning steps and adjusts local weights through clause comparisons. The method uses abstract syntax trees to compare student-generated clauses with the reference SQL and other execution-correct samples, accounting for their query context. It then propagates SQL weights to the corresponding reasoning steps through aligned SQL fragments. These weights scale the existing teacher feedback without changing the task reward or adding inference-time procedures. Trained only on BIRD with a Qwen3.5-4B student, our method achieves 64.64% execution accuracy on BIRD-dev, improving over vanilla OPD by 3.58 percentage points. It also improves execution accuracy across five Spider benchmarks by 1.75–2.61 points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.