TraceAhead: Learning Non-Myopic Control for LLM Search Agents from Branched Executions
Abstract
LLM search agents often need to decide whether to keep searching and, if so, which direction to pursue, even though the value of a search action may only become apparent several steps later. Judging a search action only by its immediate effect can therefore overlook actions that enable valuable future searches. We introduce TraceAhead, a lightweight value-based controller for non-myopic control of LLM search agents. TraceAhead models search control as a finite-horizon decision problem, with stopping as the zero-value reference and each search action valued by its immediate net gain together with the continuation value it enables. It learns these history-dependent action values from branched training executions: starting from shared checkpoints, alternative actions are executed and regularized backward fitting propagates downstream value from successor histories to earlier decisions. At inference time, the learned values guide the agent's decision of whether to stop and, if continuing, which search direction to pursue, while only the selected action is executed along a single search path. We evaluate TraceAhead on three multi-hop question answering benchmarks with three LLM search actors. At development-selected operating points with similar token use, TraceAhead achieves higher final-answer quality than three alternative search-control policies in all nine model–benchmark settings, with gains of up to 8.2 F1 points over the immediate-value policy. Additional analyses indicate that the learned continuation signal contains predictive information about future search opportunities. Together, these results show that training-time branches can provide supervision for history-dependent, non-myopic search decisions in LLM agents without requiring branch expansion at inference time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.