Perform better and spend less: Variable-compute policies in runtime-aware RL
Abstract
We study control problems where an agent is penalized for the time it takes to compute an action. Consequently, in scenarios where taking more time to compute an action results in a better outcome, a trade-off emerges: how should an agent balance the short-term cost incurred for calculating an action against the long-term total reward that taking that action would result in. We formalize such problems as runtime aware MDPs (RAMDPs) that track the compute spent by an agent for determining each action it takes, and discuss some interesting consequences of this formalization. Notably, in some cases, only variable-compute policies, i.e., policies that vary their computation time depending on the environment's state, are optimal in RAMDPs. We also conduct experiments on three simulated environments to illustrate how variable-compute policies can maintain performance while spending less compute when compared to uniform-compute policies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.