acceptodds
Under review as a conference paper at ICLR 2027

TACO: Reinforcing Multi-Turn Table Agents with Hierarchical Credit Assignment

Abstract

Table agents answer questions over multiple turns, where each turn writes code, executes it, and reads the result. Reinforcement learning on the final answer gives every decision the same credit. Credit assignment promises finer supervision, yet existing methods bring limited gains on table agents. We show that credit should be assigned hierarchically, first across turns and then within each turn, because a turn often chains several table operations of which only some are incorrect. Existing token-level methods, however, hardly distinguish the contributions of different operations within a turn. We further find that hindsight, the information revealed after a turn, shifts the probabilities of its tokens differently for correct and incorrect operations. We therefore propose TaCo, a hierarchical Table Credit Optimization framework for multi-turn table agents. Across turns, TaCo compares each turn with turns that received the same observations. Within a turn, a hindsight teacher, a periodically synchronized copy of the policy, rescores the turn with its execution feedback and a correct trajectory. The turn credit sets the direction of the update, and the probability gap between the teacher and the policy sets its size for each code and answer token. Extensive experiments on four table reasoning benchmarks across three model scales demonstrate that TaCo consistently improves multi-turn table agents and generalizes to benchmarks outside its training data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.