Beyond Latent Similarity: Credit Assignment with State Transitions for LLM Agents
Abstract
In agentic RL, feedback based on final episode outcomes does not directly distinguish the contributions of individual steps. Recent methods derive local credit from explicit interaction records or policy hidden states. For methods that share returns within latent neighborhoods formed by hidden-state similarity, pooling records by action can assign identical credit to steps with different observed transitions. To address this limitation, we introduce CAST, a method for step-level credit assignment motivated by state abstraction theory. CAST forms candidate neighborhoods using policy hidden-state similarity and selects references that match the target step in both action and observed transition, while retaining records of other actions for local comparison. It estimates step-level credit from the retained records without additional environment interaction or a learned value model. Experiments with Qwen3.5-4B and Qwen3.5-9B across four interactive environments show that CAST outperforms GiGPO and BiPACE in all eight settings, with success rate gains of up to 8.9 and 22.0 percentage points, respectively. Additional experiments extend these gains to Llama in two environments. Evaluations on 11 general capability benchmarks further show that both Qwen models retain overall performance after training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.