FDR-Controlled Selection in Reinforcement Learning via Distributional Likelihood Ratios
Abstract
In high-stakes reinforcement learning, reliable deployment requires identifying instances where a target policy is safe or effective, rather than relying solely on average performance. We study the problem of selecting a subset of initial states whose returns exceed a given threshold, while controlling the false discovery rate (FDR) of erroneous selections. This problem is challenging due to unobservable long-horizon returns and distribution shift in off-policy settings. To address this, we propose a unified framework for FDR-controlled selection in RL based on distributional modeling and likelihood-ratio principles. Our approach estimates the return distribution under the target policy and constructs selection rules that directly quantify evidence for desirable outcomes, avoiding the need for explicit p-value construction. We further incorporate a temporal decomposition strategy and an offline calibration procedure to enable reliable FDR control under partial observability and policy mismatch. We provide theoretical guarantees and demonstrate empirically that our method achieves accurate FDR control with improved selection power across a range of environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.