acceptodds
Under review as a conference paper at ICLR 2027

AdaGamma: State-Dependent Discounting for Temporal Adaptation in Reinforcement Learning

Abstract

The discount factor in reinforcement learning controls both the effective planning horizon and the strength of bootstrapping, yet most deep RL methods use a single fixed value across all states. State-dependent discounting is a natural way to adapt temporal credit assignment to local state structure, but naively learning a discount function inside deep actor–critic methods creates a self-referential target: the learned discount scales the same bootstrap term against which it is trained. Consequently, TD-error-based objectives can admit a degenerate shortcut, reducing target dependence on downstream value estimates instead of learning meaningful temporal structure. We propose AdaGamma, a practical framework for state-dependent discounting in deep actor–critic reinforcement learning. AdaGamma learns a bounded discount function \(\gamma_\phi(s)\) and trains it using a return-consistency objective that matches a one-step adaptive bootstrap to a multi-step return computed with a lagged, stop-gradient reference discount. This removes the direct target manipulation path and encourages the discount to encode state-dependent value propagation. We analyze the Bellman operator induced by a fixed state-dependent discount in finite tabular MDPs, proving contraction, exact soft policy iteration, and Lipschitz sensitivity under standard boundedness assumptions. Empirically, AdaGamma integrates into SAC and PPO, avoids TD-collapse, improves continuous-control benchmarks, and achieves statistically significant gains in a four-week online A/B test on the JD Logistics platform.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.