acceptodds
Under review as a conference paper at ICLR 2027

Time and Cause: Reliability-Adaptive Hybrid Credit Assignment in Reinforcement Learning

Abstract

Assigning credit to decisions with delayed consequences remains a central challenge in reinforcement learning. Causal credit assignment can reduce policy-gradient variance, but its benefits depend on accurately learned underlying causal mechanisms. In this paper, we study two challenges of this approach: unreliable learned contributions can introduce persistent policy-gradient bias, while ignoring the delay between actions and outcomes can make the relevant causal relationships more difficult to learn. We introduce ybrid elay-aware redit ssignment (HDCA), which addresses both challenges by (i) adaptively interpolating between unbiased temporal and learned causal policy-gradient estimators based on uncertainty in the estimated causal mechanisms and (ii) conditioning credit on the time gap between an action and its outcome. We support these design choices with theoretical analyses that characterize their benefits. Across a diverse suite of existing and newly designed credit-assignment benchmarks, HDCA achieves the best or tied-best final performance while improving sample efficiency and training stability. Mechanistic evaluations of gradient quality and credit-assignment accuracy provide evidence for the sources of these improvements. Finally, targeted experiments empirically validate the theoretical analyses underlying HDCA.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.