DAET: Distributional Action Equivalence Testing Integrating Aleatoric and Epistemic Uncertainty
Abstract
Modern reinforcement learning (RL) algorithms typically treat actions as distinct during learning, even when different actions lead to nearly identical long-term consequences. This action redundancy unnecessarily inflates the effective exploration space and severely degrades exploration efficiency. This motivates the need for a principled criterion for identifying actions that can be safely treated as equivalent. However, existing equivalence criteria are inadequate for this purpose: transition-based equivalence is overly restrictive, while mean-based equivalence is too permissive and insensitive to risk. More importantly, both paradigms treat action equivalence as a property of directly accessible ground truth, whereas in online RL, it must be inferred through learned estimates from finite data. To bridge this gap, we introduce Distributional Action Equivalence Testing (DAET), a novel formulation that casts action equivalence as a statistical inference problem over return distributions. DAET jointly accounts for aleatoric uncertainty in long-term outcomes and epistemic uncertainty in their learned estimates, requiring sufficient statistical evidence before equivalence is accepted. Building on this formulation, we propose DAET-Mask, an uncertainty-aware framework that operationalizes DAET for online RL. It integrates a distributional RL backbone with Deep Ensembles, a dual-gating mechanism, and an uncertainty-guided ROPE decision rule to realize safe dynamic action masking without relying on hand-crafted rules. Theoretically, we show that distributional action masking is near-optimal in the ideal case and remains robust under bounded estimation error. Experiments in both controlled and realistic environments demonstrate that DAET-Mask reliably identifies and masks redundant actions, thereby improving exploration efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.