HOPE: Hands-Off Counterfactual Policy Explanations for Reinforcement Learning Failures
Abstract
Explaining a reinforcement learning (RL) failure requires identifying when and how the agent could have acted differently to avoid failure while still completing its task. Existing counterfactual approaches modify observations, generate open-loop action sequences, or learn alternative policies, leaving a gap between explaining a specific failure and preserving the deployed policy's feedback. We present HOPE (Hands-Off counterfactual Policy Explanation), a unified framework for generating trajectory-specific counterfactual policies across continuous, discrete, and hybrid action spaces. Each counterfactual policy applies sparse, time-indexed action interventions to change the failure outcome while preserving the deployed policy's feedback at all other times. HOPE separates validity search from intervention reduction, using a graded reach-avoid residual to guide search across continuous, discrete, and hybrid action spaces. Validity-preserving deletion, refitting, and magnitude contraction then reduce intervention burden, while Pareto filtering retains nondominated trade-offs between intervention number and average magnitude. Experiments on 600 failure incidents across three navigation and driving benchmarks yield validity rates of 81%, 74%, and 83% on the common solvable subsets. Across all benchmarks, HOPE ranks first or second in validity and first in single-intervention rate among the evaluated RL counterfactual methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.