DriftBench: An Execution-Based Benchmark for LLM SQL Query Repair Under Schema Evolution
Abstract
Production analytics break during schema evolution: a benign table or column rename can invalidate thousands of downstream queries. Large language models (LLMs) are an attractive repair tool. However, the field lacks execution-based benchmarks that (i) guarantee the drift actually affects the query, (ii) preserve semantics through migrated data, and (iii) score correctness beyond syntactic validity. We present , a procedural generator and evaluation harness for SQL repair under schema evolution. It builds paired original and drifted databases in isolated namespaces, migrates data through canonical migration plans, and derives oracle repairs by AST rewriting, so a repair counts as correct only if it reproduces the pre-drift result set. Across 18,624 system evaluations spanning seven LLMs, a deterministic schema-diff oracle, and a fuzzy heuristic, the best model reaches 92.0% semantic equivalence on the 2,000-instance rename benchmark against 13.7% for the heuristic. Under the default prompt, execution success overstates repair quality for every LLM on every rename suite; on the 292-instance rename suite, the gap is 5.9–19.9 points for LLMs and 53.1 for the heuristic; given both schema versions, a deterministic diff solves renames perfectly, yet on structural drift it collapses to 0–6.7% while LLMs reach 100%; and on column semantic transforms, six of nine systems execute 98–100% of the time while only two exceed 20% equivalence. To test transfer, we mine the histories of 40 open-source applications, identify 855 schema-verified renames across 34 of them, and build a 300-instance real-migration slice. There, seven current models execute 94–100% of the time but are correct on only 75–90%, and 93% of their failures are queries that run and return wrong rows; the real migration's rename list lifts them to 93–100%. Model rankings on the synthetic held-out-vocabulary and cross-domain suites predict real-slice rankings (Spearman and ), and the semantic-transform gap persists in current models (46–87% equivalence at 100% execution).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.