acceptodds
Under review as a conference paper at ICLR 2027

Controlling the Overestimation Bias of Q-learning From an Action-Selection and Action-Evaluation Overlap Perspective

Abstract

The well-known estimation bias causes Q-value estimates to be overestimated or underestimated, potentially impairing learning and policy performance. This bias is closely related to the sharing of Q-functions between action selection and value evaluation. However, existing methods lack a systematic analysis of how the sharing degree affects estimation bias, providing limited theoretical guidance for bias minimization across different settings. We formalize the sharing degree as the overlap between the Q-functions used for action selection and value evaluation, and analyze how this overlap affects estimation bias and should be chosen to minimize bias. Under the stated assumptions, we prove that the signed estimation bias is a monotonic function of this overlap, and that the bias-minimizing overlap is governed by the signal-to-noise ratio (SNR). Motivated by this finding, we propose Selective-Overlap Q-learning (SOQ). SOQ uses all Q-functions for value estimation and selects a Q-function subset for action selection according to a state-dependent proxy for SNR, aiming to approximate the bias-minimizing overlap and reduce estimation bias. By adapting the overlap between action selection and value evaluation, SOQ improves the reliability of Q-value estimation across different settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.