Show or Stow: Learning Evidence Admission from Stage-Specific Utility in Multi-Agent Search
Abstract
Agentic search has become a central approach to reasoning-intensive information retrieval. When multiple agents carry it out, the evidence they collect guides further search and supports final-answer generation, but the same evidence can have different utility at these two stages. Existing approaches organize shared evidence or learn what to share without separately learning its value for each use. To mitigate this mismatch between stages, we propose StageAdmit, an evidence admission framework that keeps every collected item in an evidence pool and regresses stage-specific keep-mask utilities to learn separate admission policies for continued search and final-answer generation. A predictor with stage-specific heads decides which evidence enters each search agent's panel during search and which enters the writer's panel for the final answer. The predictor is trained on paired comparisons that keep or mask a piece of evidence, measuring its effect on final-answer quality and subsequent tool calls for continued search, and its marginal gain for final-answer generation. Because all evidence remains in the pool, evidence withheld from search agents can still reach the writer. Experiments on WideSearch and GISA show that StageAdmit outperforms all evaluated baselines, surpassing the strongest baseline by 1.5 and 1.6 Item-F1 points with GLM-5, while reducing completion time by 13.1% and writer input tokens by 36.4%. The same frozen predictor also improves GPT-5.1 and Qwen3.8-27B agents without retraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.