acceptodds
Under review as a conference paper at ICLR 2027

Perform better and spend less: Variable-compute policies in runtime-aware RL

Abstract

We study control problems where an agent is penalized for the time it takes to compute an action. Consequently, in scenarios where taking more time to compute an action results in a better outcome, a trade-off emerges: how should an agent balance the short-term cost incurred for calculating an action against the long-term total reward that taking that action would result in. We formalize such problems as runtime aware MDPs (RAMDPs) that track the compute spent by an agent for determining each action it takes, and discuss some interesting consequences of this formalization. Notably, in some cases, only variable-compute policies, i.e., policies that vary their computation time depending on the environment's state, are optimal in RAMDPs. We also conduct experiments on three simulated environments to illustrate how variable-compute policies can maintain performance while spending less compute when compared to uniform-compute policies.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.