acceptodds
Under review as a conference paper at ICLR 2027

DAET: Distributional Action Equivalence Testing Integrating Aleatoric and Epistemic Uncertainty

Abstract

Modern reinforcement learning (RL) algorithms typically treat actions as distinct during learning, even when different actions lead to nearly identical long-term consequences. This action redundancy unnecessarily inflates the effective exploration space and severely degrades exploration efficiency. This motivates the need for a principled criterion for identifying actions that can be safely treated as equivalent. However, existing equivalence criteria are inadequate for this purpose: transition-based equivalence is overly restrictive, while mean-based equivalence is too permissive and insensitive to risk. More importantly, both paradigms treat action equivalence as a property of directly accessible ground truth, whereas in online RL, it must be inferred through learned estimates from finite data. To bridge this gap, we introduce Distributional Action Equivalence Testing (DAET), a novel formulation that casts action equivalence as a statistical inference problem over return distributions. DAET jointly accounts for aleatoric uncertainty in long-term outcomes and epistemic uncertainty in their learned estimates, requiring sufficient statistical evidence before equivalence is accepted. Building on this formulation, we propose DAET-Mask, an uncertainty-aware framework that operationalizes DAET for online RL. It integrates a distributional RL backbone with Deep Ensembles, a dual-gating mechanism, and an uncertainty-guided ROPE decision rule to realize safe dynamic action masking without relying on hand-crafted rules. Theoretically, we show that distributional action masking is near-optimal in the ideal case and remains robust under bounded estimation error. Experiments in both controlled and realistic environments demonstrate that DAET-Mask reliably identifies and masks redundant actions, thereby improving exploration efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.