Table-Zero: Self-Improvement for Complex Table Reasoning without Additional Human Supervision
Abstract
Complex table reasoning requires agents to combine evidence across multiple tables through multi-step reasoning, but high-quality training data for these tasks is costly to build. Self-improving reinforcement learning (RL) offers a promising alternative: a model generates its own training questions and learns from them without human labels. However, existing self-improving methods struggle in this setting: it is hard for a model to generate questions that require multiple tables, and the reference answers to such questions may be wrong. To address this, we introduce \method, a self-improvement framework that uses the table collection itself to choose which tables each question combines and to compute its answer. \method selects related tables through two intrinsic properties of tables, joinability and decomposability. A proposer writes a question and reference SQL over these tables, and executing the SQL supplies the reference answer. The proposer and the solver then improve in turn: the proposer is trained with DPO on its own checked questions, and the solver is trained with RL on the resulting tasks, for three iterations. Across seven datasets, \method achieves an average exact match of 32.3% with Qwen3-8B and 37.0% with Qwen3-14B, compared with 20.4% and 24.3% for , respectively. At both model sizes, it exceeds the average accuracy of adapted self-improvement baselines and RL trained on human-annotated tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.