acceptodds
Under review as a conference paper at ICLR 2027

Stop or Continue? Counterfactual Timing Supervision for Search Agents

Abstract

Multi-turn search agents must learn not only what to search for, but also when further interaction is worth its cost. Poor timing leads to two opposite errors: stopping prematurely with incomplete evidence, or continuing to search after sufficient evidence has already been acquired. Direct supervision of this timing decision requires a counterfactual comparison between the utility of continuing and that of stopping from the same decision context, whereas a standard rollout reveals only the outcome of the chosen action. We introduce Action-Timing, a cost-aware reinforcement learning framework with an explicit STOP/CONTINUE timing decision. Using main-trajectory branches together with matched same-action and counterfactual rollouts organized in a knowledge-state graph, Action-Timing estimates the Marginal Utility of Continuation (MUC), defined as the difference between the expected cost-adjusted downstream utilities of continuing and stopping at that decision point. MUC directly supervises whether further interaction is worthwhile, while continuation plans receive credit relative to the same stopping reference, separating learning when to continue from learning how to continue. Experiments on search-agent benchmarks show improved task performance over competing methods. Trajectory analyses further illustrate bidirectional timing correction, showing that Action-Timing can continue beyond premature stopping points to recover correct answers while terminating redundant searches earlier without sacrificing answer correctness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.