acceptodds
Under review as a conference paper at ICLR 2027

Useful empowerment approximation for its use in Learning Agents

Abstract

Process (closed-loop) empowerment is a relevant measure for agency, with promising uses for intrinsic motivation and reinforcement learning. However, its prohibitive computational cost, especially in discrete environments, precludes many of its applications to anything but toy examples and the smallest environments. We introduce a reformulation of process empowerment for Markovian dynamics under which the empowerment computation becomes tractable, and derive from it a variant of the Alternating Blahut-Arimoto algorithm whose per-iteration cost is polynomial in the number of states, actions and horizon compared to previous exponential in the horizon per-iteration cost; making empowerment a viable signal in more complex environments. Moreover, we introduce even faster approximations of empowerment via early-stopping, showing that few iterations are enough to compute empowerment to a reasonable precision, and a warm-start initialisation reusing other states' empowerment information to hasten computation. Across six Minigrid environments, the resulting approximations reach a STARC distance, a reward comparison metric, of 0.025 from exact empowerment making the approximation behaviourally interchangeable with the exact signal at almost no cost. Our approach achieved this in between 1.0 and 6.4 iterations per state on average compared to between 75 and 1275 for the exact computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.