acceptodds
Under review as a conference paper at ICLR 2027

SQL-Structure-Aware On-Policy Distillation for Text-to-SQL

Abstract

On-policy distillation (OPD) provides token-level teacher feedback for training Text-to-SQL models. However, uniform weighting allows lengthy reasoning text to dominate the distillation objective and does not account for how tokens jointly express SQL operations, columns, and values. We propose a SQL-structure-aware weighting method for OPD that controls the total weight assigned to final SQL and reasoning steps and adjusts local weights through clause comparisons. The method uses abstract syntax trees to compare student-generated clauses with the reference SQL and other execution-correct samples, accounting for their query context. It then propagates SQL weights to the corresponding reasoning steps through aligned SQL fragments. These weights scale the existing teacher feedback without changing the task reward or adding inference-time procedures. Trained only on BIRD with a Qwen3.5-4B student, our method achieves 64.64% execution accuracy on BIRD-dev, improving over vanilla OPD by 3.58 percentage points. It also improves execution accuracy across five Spider benchmarks by 1.75–2.61 points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.