acceptodds
Under review as a conference paper at ICLR 2027

Learning Task-State Abstractions for Long-Horizon Database Agents

Abstract

Tool-using database agents must track what they know, what they have tried, and which evidence should guide the next action. We study how structured task-state representations support this sequential decision problem. TSADB combines schema topology, query intent, execution traces, and error semantics into a learned state consumed by a lightweight PPO policy; a frozen language model generates tool arguments. On BIRD-SQL and Spider 2.0-Lite, the system improves execution accuracy by 10.0 and 10.4 percentage points over the strongest matched baseline in each main comparison. A separate training-budget-matched comparison with RL-ReAct yields a 7.9-point BIRD gain. A state-by-reward factorial shows complementary contributions: flat-history encoding reduces accuracy by 9.1 points with the progress reward, while removing that reward costs 10.5 points with the learned state. The state effect grows to 18.5 points on tasks requiring at least 16 calls. Across policy classes, supervised control also benefits from the representation, and PPO adds 13.1 points over supervised control on long tasks in a shared 800-task evaluation. The full system retains 74% of its accuracy under combined schema perturbations. Experiments further examine official conversational interaction, transfer with limited error-projection adaptation, termination, and reward construction. These results support learning compact, structured states for long-horizon database control.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.