acceptodds
Under review as a conference paper at ICLR 2027

DisTAR: Disentangled Token Aggregation for Reasoning and Action in LLM Agent Training

Abstract

Reinforcement Learning (RL) has been pivotal in empowering autonomous agents to solve long-horizon, multi-turn tasks. However, most existing methods directly inherit standard token aggregation paradigms in single-turn reasoning, thereby overlooking the fundamental divergence between reasoning and action: reasoning requires thorough exploration within a single turn, whereas action demands concise and precise sequences to prevent error propagation. In this paper, we first identify and formalize this “aggregation dilemma" in agentic RL: applying token-mean to action tokens inadvertently assigns higher gradient weights to longer, redundant action chains, while enforcing sequence-mean on reasoning tokens hinders the model's Chain-of-Thought (CoT) exploration. To resolve this conflict, we propose DisTAR (Disentangled Token Aggregation for Reasoning and Action), a novel framework that decouples the token aggregation paradigm. Specifically, DisTAR enforces a token-mean constraint on reasoning tokens to encourage sufficient reasoning before action, while executing a mathematically equivalent sequence-mean reweighting mechanism on action tokens to optimize for concise action chains. Furthermore, at the trajectory level, an inverse-step penalty is integrated into the advantage estimation to proactively suppress redundant paths. Finally, we introduce a globally normalized surrogate loss to maintain stable training convergence under this heterogeneous token aggregation strategy. Extensive experiments on WebShop, AlfWorld, and SearchQA across multiple foundation models (e.g., Qwen-2.5, Qwen-3) demonstrate that DisTAR significantly boosts agent performance. Notably, on WebShop with Qwen-3-4B, DisTAR enhances the success rate of the baseline GRPO from 68.6% to 82.8%, while consistently inducing shorter, more precise action chains alongside richer CoT. Code is available at https://anonymous.4open.science/r/DisTAR-A387.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.