acceptodds
Under review as a conference paper at ICLR 2027

Show or Stow: Learning Evidence Admission from Stage-Specific Utility in Multi-Agent Search

Abstract

Agentic search has become a central approach to reasoning-intensive information retrieval. When multiple agents carry it out, the evidence they collect guides further search and supports final-answer generation, but the same evidence can have different utility at these two stages. Existing approaches organize shared evidence or learn what to share without separately learning its value for each use. To mitigate this mismatch between stages, we propose StageAdmit, an evidence admission framework that keeps every collected item in an evidence pool and regresses stage-specific keep-mask utilities to learn separate admission policies for continued search and final-answer generation. A predictor with stage-specific heads decides which evidence enters each search agent's panel during search and which enters the writer's panel for the final answer. The predictor is trained on paired comparisons that keep or mask a piece of evidence, measuring its effect on final-answer quality and subsequent tool calls for continued search, and its marginal gain for final-answer generation. Because all evidence remains in the pool, evidence withheld from search agents can still reach the writer. Experiments on WideSearch and GISA show that StageAdmit outperforms all evaluated baselines, surpassing the strongest baseline by 1.5 and 1.6 Item-F1 points with GLM-5, while reducing completion time by 13.1% and writer input tokens by 36.4%. The same frozen predictor also improves GPT-5.1 and Qwen3.8-27B agents without retraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.