The Truth Is Out There: Separating Owner Definitions from Database Facts in Text-to-SQL
Abstract
Text-to-SQL benchmarks give agents two kinds of missing information at once. The database settles some of it, such as which column holds a concept or how its values are encoded. The data's owner decided the rest, such as where a threshold lies or which categories count, and no query recovers it. Because benchmark hints supply both together, a score cannot say which one an agent lacked. We separate them with definition relocation. For an underspecified question, we write several definitions an analyst could hold, make one of them correct, and move it between the prompt, a glossary table in the database, the schema text, the agent's transcript, and the data owner. Each definition implies a different answer, so the answer shows which one the agent acted on. Where the definition sits matters as much as whether it exists. Open-weight models almost never query the glossary, so a definition stored there is worth nothing to them, while the same rows returned by a query run on their behalf close 35 to 85% of the gap to stating the definition in the prompt. Agents that explore on their own do recover stored definitions, one sentence of instruction makes a hosted model do the same, and how often an agent queries the glossary predicts what a stored definition is worth to it. Documentation of what the database already settles behaves differently: it helps small models and adds nothing measurable for strong ones. A second set of concepts reproduces these results.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.