acceptodds
Under review as a conference paper at ICLR 2027

DARES: Decision-Aligned Process Supervision for Multi-Turn Text-to-SQL Agents

Abstract

Multi-turn Text-to-SQL agents must translate evolving user intents and database feedback into a sequence of tool-use and query decisions. Terminal execution rewards indicate whether a trajectory succeeds but not which decisions were responsible. Step-level rewards offer finer credit, yet their predefined step boundaries need not align with the semantic decisions that determine SQL correctness: one step may mix correct and faulty decisions, and one decision may span several steps. We propose DARES, which automatically discovers decision-aligned reward units from policy rollouts. Its Automatic Reward Unit Discovery (ARUD) constructs localized SQL contrasts, each pairing a correct and a faulty realization of one decision, verifies through controlled execution that the correct side indeed yields better outcomes, and maps the verified decisions back to the policy-generated token spans that express them. A decision-aligned process reward model learns from these contrasts and provides localized credit during GRPO, while execution rewards retain trajectory-level supervision. On interactive Text-to-SQL, DARES improves normalized reward over execution-only GRPO by up to points. The result show that decision-aligned process supervision improves policy optimization across diverse multi-turn database tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.