WHEN AND HOW DOES A REINFORCEMENT LEARNING OPTIMIZER HELP OUT-OF-DISTRIBUTION DETECTION? A THEORETICAL STUDY
Abstract
Out-of-distribution (OOD) detection in dynamic open-world environments requires a model to continually adapt to evolving data distributions while generalizing to covariate-shifted inputs and rejecting semantic-shifted OOD examples. Most existing OOD detection methods optimize only the current-step objective and do not explicitly account for how post-deployment environment changes affect future OOD behavior. In this paper, we establish a theoretical framework that explains how and when reinforcement learning (RL)-guided update improves dynamic OOD detection. Our RL-guided optimizer explicitly favors updates that reduce the semantic OOD false positive rate over time. We analyze how the value function shapes the updates and affects the training trajectory, and derive a temporal error decomposition that separates the change in evaluation risk into a model-change term, controlled by the optimizer, and an environment-change term, derived by distribution shift. This decomposition yields measurable conditions under which the RL-guided optimizer lowers the semantic-OOD false positive rate relative to GD without degrading classification generalization error. Experimental evidence support better performance of the RL-guided optimizer over a GD optimizer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.