acceptodds
Under review as a conference paper at ICLR 2027

Budget-Aware State Criticality via Optimal Interventions under Decaying Agency

Abstract

In reinforcement learning, _critical states_ are situations where the chosen action has a disproportionately large impact on the final outcome. Identifying such critical states is a fundamental challenge, with applications in explainability, credit assignment, and human oversight of autonomous agents. We present CODA (**C**riticality via **O**ptimal Interventions under **D**ecaying **A**gency), a framework in which an agent can intervene by overriding the action of a fallback policy, but each intervention can make future ones less likely to take effect. Critical states are those where the agent still chooses to intervene, making criticality explicitly relative to the fallback policy. We show that the resulting problem decomposes into a cascade of coupled optimal stopping problems, and build on this structure to develop algorithms for both tabular and continuous environments. We prove that critical states exist if and only if the fallback policy is suboptimal, and that, unlike in penalty-based approaches, they are invariant to the scale of the reward. Empirically, CODA identifies intuitive critical states in grid worlds and continuous control, and, when used to allocate a limited intervention budget, matches or exceeds the returns achieved with State Importance and Lazy-MDPs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.