acceptodds
Under review as a conference paper at ICLR 2027

Solve More with Fewer Steps: Assigning Credit by Uncovering Dependencies in Agentic RL

Abstract

Agentic reinforcement learning (RL) is a powerful technique for improving the multi-step reasoning of agents. The success of agentic RL largely relies on whether it can properly assign credit for final outcomes to intermediate steps, namely the credit assignment problem (CAP). However, solving the CAP remains highly challenging, as long reasoning trajectories often obscure the contribution of each step. To tackle this problem, we propose **CreditFlow**, a novel credit assignment method that directs agents to focus on critical steps, thus improving both reasoning performance and efficiency. Our key idea is to assign credit based on *underlying dependencies* (whether step depends on the result of step ), rather than the *temporal order* (whether step occurs after step ). The major challenge in developing CreditFlow is to efficiently distinguish underlying dependencies from the temporal adjacency in raw trajectories. To uncover the dependencies, CreditFlow introduces a rule-based provenance tracking mechanism to dynamically trace how later steps use results from earlier ones, *incurring negligible overhead* and *avoiding LLM-induced hallucinations*. Experiments on eight mainstream benchmarks show that CreditFlow outperforms the strongest baseline by **5 points** on average, while reducing the average number of inference steps by over **20%**.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.