Uncovering Latent Reasoning from Action Trajectories for Long-Horizon Agents
Abstract
Training Large Language Model (LLM) agents for long-horizon tasks requires high-quality trajectory data that tightly interleaves reasoning and action. While action data is abundant in environment logs (e.g., human interaction traces in computer-use settings), the underlying reasoning is inherently latent and difficult to collect. In this paper, we introduce PAIR, a framework designed to synthesize complete training trajectories by uncovering latent reasoning directly from observable action sequences. The key idea is to train a reasoning generator via reinforcement learning through a novel reward mechanism: an action compatibility reward that maximizes alignment between the synthesized reasoning and subsequent ground-truth actions, paired with a distribution consistency reward that constrains the reasoning to remain within the target policy's natural distribution, preventing trivial shortcuts. Across diverse domains, including deep research, computer use, and web navigation, PAIR consistently outperforms action behavior cloning and state-of-the-art reasoning reconstruction techniques. Furthermore, PAIR bootstraps agent performance on unfamiliar tasks and serves as a strong initialization for subsequent agentic RL, demonstrating that action-only trajectories can serve as a powerful supervisory signal for long-horizon agent training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.