Towards Adaptive Temporal Credit Assignment via Bidirectional Fractional Operators
Abstract
In continuous control and sequential decision-making tasks, agents may receive only an episodic return, with no intermediate rewards. Learning from such feedback requires connecting decisions to the terminal outcome. However, recurrent compression can weaken early contextual information, while fixed exponential weighting can attenuate distant learning signals. Moreover, history encoding and credit propagation may require different temporal scales. To address these challenges, we propose Fractional Credit Assignment (FCA), a framework based on two oppositely directed Caputo fractional operators. First, the left-sided operator introduces nonlocal history encoding through a power-law kernel. This allows early trajectory information to directly contribute to later representations. Second, the right-sided operator provides power-law temporal weights for backward credit assignment. Terminal-to-history attention combines content matching with these weights to redistribute the episodic return across steps. The fractional orders α1 and α2 are learned independently, allowing the two temporal scales to adapt to each task. We evaluate FCA under episodic-return-only feedback on seven MuJoCo tasks and four DeepMind Control Suite tasks. The latter use fixed 1,000-step episodes. FCA achieves the highest final average return on nine of the eleven tasks. On Swimmer, Ant, and Walker2d, it improves average returns over the strongest baseline in the same feedback setting by 90.7%, 53.3%, and 42.8%, respectively. It also outperforms SAC trained with dense rewards on five MuJoCo tasks. Ablation studies support the benefits of nonlocal history encoding and independently learned temporal scales for long-horizon credit assignment. The code is available at https://anonymous.4open.science/r/FCA-FBEC.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.