TABRSI: A RECURSIVE SELF-IMPROVING MULTI- AGENT SYSTEM FOR TABLE UNDERSTANDING AND REASONING
Abstract
Autonomous recursive self-improvement enables coding agents to iteratively refine their internal logic, yet whether these mechanisms generalize beyond software engineering to complex tabular reasoning, and whether performance lifts stem from algorithmic search depth or structural task levers, remains an open question. We investigate self-improving agents across seven structured-table benchmarks by formalizing a unified design space spanning search policies (greedy Darwin Gödel Machines vs. Clade-Metaproductivity in Huxley-Gödel Machines) and system architectures (monolithic solo solvers vs. modular role teams). To overcome the expressivity bottlenecks of monolithic scripts, we introduce DGM-Flow and HGM-Flow, which decompose the agent into an independently evolvable pipeline of specialized roles, Parser, Planner, Solver, Verifier, and Formatter}, governed by topological gating, code mutations, and per-role reasoning modalities (Chain-of-Thought vs. Program-of-Thought) backed by a finite-sample non-regression guarantee. Across all seven benchmarks on a unified DeepSeek-v4-flash backbone, empirical gains follow a regime law: performance lifts do not track evolutionary search depth, but rather the presence of an executable tool or structural lever targeting an addressable dominant bottleneck. Where addressable bottlenecks exist, role decomposition and solver self-improvement achieve held-out verified wins, including an absolute gain of in accuracy on HiTab via hierarchical header reconstruction, an absolute gain of on MultiHiertt through algorithmic parse-then-compute routines, and up to absolute accuracy recovery on MMTU subtasks via program-aided verification and formatting. Conversely, search-rule deltas remain within evaluation noise on interpretation-bound or saturated tasks, co-evolving solvers jointly with role teams fails to generalize over decoupled application, and semantic LLM-judge evaluations expose that unconstrained solver self-edits frequently optimize superficial formatting envelopes rather than genuine reasoning capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.