Context-Grounded Data Curation and Gate-First RL Reward Design for SQL Autocomplete
Abstract
SQL autocomplete models trained on user editing logs suffer from a supervision mismatch: the recorded completion often depends on information absent at the cursor, such as a table not yet referenced, a literal not in the session, or a clause whose intent emerges only with further typing. We address three coupled problems: what constitutes a valid training target, how to measure both suggestion quality and firing-abstention behavior, and how to train the model to balance useful completions with calibrated silence. Raw editing logs contain content the model could not have inferred at the cursor; we enforce context grounding by trimming each SFT target to its longest inferable prefix and retaining empty targets as silence labels, raising precision by nearly 13 percentage points. To further improve suggestion quality and calibrate the firing decision, we apply RL with deterministic abstention and grounding gates before a graded LLM-judge reward, which prevents the silent collapse observed with a judge-only reward and raises pass@1 on predictable triggers from 35.8% to 43.4%. Since public benchmarks do not score silence, we reconstruct intended completions from real editing sessions and design a selective evaluation framework that jointly measures suggestion quality and firing-abstention behavior. In matched seven-day deployment windows, the 4B dense model, trained on no customer data, raises live acceptance from 17.8% to 26.35% over the previous 30B-A3B MoE while reducing median latency by 71%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.