acceptodds
Under review as a conference paper at ICLR 2027

From Trajectory Evidence to Control: Latent Control-State Inference for Long-Horizon Language Agents

Abstract

Language agents often fail in long-horizon tasks despite making locally reasonable decisions. A recurring difficulty is that accumulated trajectory evidence may call for a change in how the agent should proceed, while next-action generation continues along the current course. We formulate this missing intermediate decision as latent control-state inference and introduce an Explore–Exploit–Recover () controller. At each step, infers a trajectory-conditioned control state consisting of a control mode and a trajectory-specific target, and uses it to regulate the proposed action: Explore resolves missing task-relevant evidence, Exploit advances a supported course of action, and Recover breaks from one rendered ineffective by feedback. We evaluate on ALFWorld, WebArena-Lite, and ToolSandbox. Across environments, consistently improves end-to-end task performance and performance accumulated over the interaction budget. Behavioral analysis shows that acts as a sparse, mode-conditioned control layer, preserving most proposals while concentrating explicit corrections under particular inferred control states. Controlled interventions further indicate that both control-state content and its trajectory-aligned enactment contribute to the observed gains. These results suggest that effective long-horizon agency benefits from an explicit control layer between trajectory understanding and next-action generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.