HarnessSQL: HARNESS-NATIVE TRAINING FOR SQL AGENTS IN REALISTIC DATABASE ENVIRONMENTS
Abstract
Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases—inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses. This creates a critical train–deploy mismatch, as the execution harness that mediates this interaction is introduced only at inference time. To bridge this gap, we propose **HarnessSQL**, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning. **HarnessSQL** builds isolated, executable database environments paired with hidden execution oracles, rolls out teachers directly inside the target SQL harness, and retains only verified trajectories for full-sequence SFT, followed by execution-reward RL. Across Spider 2.0-SQLite, **HarnessSQL** dramatically boosts the execution accuracy of compact models, raising Qwen3-8B from 15.5% to 45.2% and Qwen3-14B from 22.2% to 54.8%, while transferring effectively to out-of-distribution interactive benchmarks such as BIRD-Interact and LiveSQLBench. Our findings demonstrate that training database agents directly within their execution harness is essential for mastering complex, long-horizon database workflows.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.