acceptodds
Under review as a conference paper at ICLR 2027

Policy Gradients for Average Path-Dependent Utility in Continuing MDPs

Abstract

Reinforcement learning typically represents desired behavior through additive per-step rewards, which can be difficult to design, and may fail to capture preferences that depend on more complicated patterns or entire trajectories. Unlike MORL, AverageRL, or General Utility RL, we study average path-dependent utility (APDU), a long-run average objective for continuing tasks in which performance is evaluated using trajectory-level utility. We restrict our attention to ergodic Markov decision processes and parametrized Markovian policies. For this setting, we first consider the case where the utility functional is known. We show that the APDU objective admits a policy-dependent proxy-reward whose expectation under the stationary state–action distribution exactly recovers the original objective. This proxy reward enables us to define a differential action-value function and derive an average-utility policy-gradient theorem. We then address the more challenging (and versatile) setting in which the utility functional is unknown and only trajectory-level utility feedback is observable. We introduce an action-value quantity that can be estimated directly from such feedback, establish its asymptotic relationship with the known utility action-value function, and derive a corresponding policy-gradient theorem. Building on these results, we propose Average Utility Actor–Critic (AUAC), an algorithm for optimizing APDU from trajectory-level feedback. Experiments validate our theoretical findings and demonstrate the effectiveness of AUAC on path-dependent control tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.