acceptodds
Under review as a conference paper at ICLR 2027

Code2Evidence: Explicit Knowledge Organization for Multi-Agent Search and Reasoning

Abstract

Retrieval agents that search a corpus over many steps keep everything they have done in one interaction history, where the few observations that support the answer sit among irrelevant hits, search errors, failed commands, and the agent's own reasoning traces. This mixture can mislead later search and the final reasoning step, and truncating or summarizing the history can discard supporting evidence along with the noise. We argue that retrieval agents should keep two separate states: a transient interaction history that may be compressed freely, and a persistent evidence state of source spans, with provenance, that the agent explicitly banks during search. The evidence state is never compressed, guides further search, and is the only input to the final answer. We implement this separation in Code2Evidence, a coding-based retrieval agent that searches raw corpora with shell commands. Holding the model, tools, turn budget, and compression policy fixed, answering from the evidence state matches or beats answering from the compressed history in all ten task–backbone pairs across six Wikipedia QA datasets (MHQA), FinQA, and BioProBench with GPT-5.4-mini and GPT-5.5. With GLM-5.3 on MHQA, adding the evidence state to a single agent raises meaning-match accuracy from 74.3% to 80.7%, and reconciling the evidence states of three role-specialized searchers reaches 85.2%. Code2Evidence is also more accurate than DCI-Agent-Lite on FinQA with all four backbones and on MHQA (85.2% vs. 79.4%), with the largest gains for GPT-5.4-mini on Legal RAG Bench and BioProBench.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.