acceptodds
Under review as a conference paper at ICLR 2027

Efficient Agentic Reinforcement Learning with Tool-Aware Credit Assignment

Abstract

Agentic reinforcement learning (RL) equips large language models (LLMs) with tool-use capabilities that substantially improve reasoning on complex tasks. However, incorporating external tool feedback perturbs the model's predicted token distribution, introducing distribution shifts that destabilize policy optimization. To address this issue, we propose TACT, an efficient framework that stabilizes policy learning through tool-aware credit assignment. Specifically, to quantify a tool's impact on the reasoning process, TACT performs counterfactual reasoning by masking tool observation tokens. By computing the information gain between the original trajectory and its counterfactual counterpart, we establish a rigorous token-level metric for credit allocation. Guided by this metric, TACT hierarchically refines the optimization process. At the trajectory level, executions with failed tool outcomes are filtered out to prevent erroneous feedback from exacerbating distribution shifts. At the token level for successful trajectories, the information gain serves as the criterion for shaping token-level advantages, allocating higher credit strictly to tokens where tool interaction meaningfully advances the final outcome. By aligning policy updates with validated tool contributions, TACT mitigates the error interference from failed trajectories while reinforcing the effect of tokens with higher information gain. Extensive experiments on 9 benchmarks spanning mathematical reasoning, code generation, and function calling, across 3 model scales, demonstrate the superiority of TACT over existing methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.