acceptodds
Under review as a conference paper at ICLR 2027

TDS: A Table-Driven Parallel Multi-Agent Search System with Asymmetric Credit Gating

Abstract

Agentic search answers a complex question by decomposing it into subproblems and gathering evidence over successive steps. Supervision comes from a single signal, the reference answer to the original question, while the intermediate subproblems carry no labels of their own. Learning from this supervision raises two questions. The first concerns how to organize intermediate results, which requires balancing sufficient evidence against concise records while maintaining a consistent search structure during training and inference. The second concerns assigning credit to intermediate decisions in multi-agent search, where value-based advantages depend on accurate state-value estimates across diverse search states, group-relative advantages do not distinguish steps, penalizing useful ones and reinforcing unnecessary ones, and dense supervision from reward models or LLM graders requires extra models or annotations and inherits the graders' biases. To address these two questions, we present TDS (Table-driven Dual-tier Search), a multi-agent search system that pairs a table-driven search structure with asymmetric credit gating. The structure is a global table of subproblem nodes, shared by training and inference, through which a planner and parallel executors coordinate. Each node keeps a subproblem's answer and a concise evidence summary written by an executor, so the planner can build on any completed result without carrying the search history, and independent subproblems are solved in parallel. We train the executor first, then freeze it and train the planner with the gate. During training, the planner is prompted after each expansion to answer the original question from the current table, and the first answer that matches the reference marks the steps that sufficed. Successful trajectories then reinforce only these steps rather than later unnecessary expansions, and failed trajectories no longer penalize them, which yields step-level credit without a critic or a grader. Under a common training and evaluation protocol, TDS improves seven-benchmark macro-average exact match over the strongest baseline at each scale by 2.2 percentage points (5.7% relative) with Qwen2.5-3B-Instruct and by 1.9 percentage points (4.4% relative) with Qwen2.5-7B-Instruct. Ablations confirm that both the table structure and the asymmetric gate contribute to these gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.