acceptodds
Under review as a conference paper at ICLR 2027

Dynamic Causal Tree: Interventional Credit Assignment for LLM Agent

Abstract

Reinforcement learning improves multi-turn decision making in LLM agents, but trajectory-level rewards obscure the contribution of individual decisions. Failed trajectories can suppress credit for useful decisions, while successful ones can reinforce harmful decisions. We formalize this mismatch through same-history interventions, separate causal relevance from policy-update credit, and bound how credit error distorts the expected policy update. We propose Dynamic Causal Tree (DCT), which selectively constructs shared-prefix branches to approximate exhaustive intervention comparisons. The resulting tree supports a dual readout: pairwise branch contrasts estimate causal relevance, while policy-weighted aggregation yields signed credit. Experiments on multi-hop QA across model scales show consistent gains over trajectory-level and tree-based RL baselines, while independent intervention tests support more faithful decision attribution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.