acceptodds
Under review as a conference paper at ICLR 2027

How Harnesses Reshape Agentic Search: Training Stateless Large Reasoning Models for Stateful Search

Abstract

Agentic search requires reasoning over evidence accumulated across interactions, yet the underlying language model is stateless across invocations. Existing approaches typically bridge this gap by carrying interaction histories across turns and applying post-training techniques. However, growing contexts can obscure earlier evidence and force the model to reconstruct search progress at each decision. Moreover, when supervision relies solely on final outcomes, training provides limited corrective feedback, leaving intermediate errors unaddressed and potentially reinforcing flawed reasoning or redundant actions. We introduce H-Search, a framework that addresses these challenges through the joint design of a stateful search harness and decision-level policy training. H-Search organizes multi-turn search into self-contained decisions: each model invocation conditions on the current structured state, with cross-turn dependencies preserved through explicit state transitions. This separation makes the same state-conditioned decision the basic unit of both search execution and policy supervision. To train within this structure, we develop harness-grounded on-policy self-distillation, where a frozen copy of the initial model uses privileged information about evidence coverage and termination readiness to provide dense token-level guidance along student-generated trajectories. The resulting policy operates within the same harness at inference without privileged information or teacher access. Across five open-domain benchmarks over Wikipedia, H-Search with a 27B backbone achieves 79.8% average recall, outperforming the evaluated same-backbone baselines by at least 11.2 percentage points. These results demonstrate the potential of coupling explicit search-state management with supervision at the level of individual decisions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.