acceptodds
Under review as a conference paper at ICLR 2027

Compact Shield Representations via Safety-Preserving Decision-Tree Approximation

Abstract

Shielding is a well-established technique for enforcing formal safety guarantees in reinforcement learning. However, its practical deployment remains limited in part by the lack of interpretability of synthesized shields. Although recent work has explored decision trees (DTs) as explainable representations of shields, exact DT representations can become prohibitively large. We address this limitation by introducing an algorithm that learns compact DT approximations of safety shields by explicitly trading permissiveness for representation complexity. To preserve the safety guarantees of the original shield, the approximation never classifies an action deemed unsafe by the original shield as safe. To mitigate the resulting performance degradation, we introduce an approximation cost that accounts for the agent's state visitation frequencies and action preferences, together with the risk induced by each action. A user-defined threshold controls the trade-off between fidelity to the original shield and DT complexity. We evaluate our algorithm across diverse domains and show that allowing a small amount of conservative approximation can substantially reduce DT size while preserving the safety guarantees of the original shield and causing at most small changes in agent performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.