Distributionally Robust Decision-Making under Reject Inference Context
Abstract
Reject inference arises when outcomes are observed only after acceptance decisions, yielding selectively labeled data that are shaped by historical policies. While existing methods mainly improve nominal learning under such one-sided feedback, they are often vulnerable to deployment shifts that alter the true conditional success probabilities. We therefore propose a distributionally robust framework for policy learning under the reject inference context. The objective is to find a robust policy that maximizes the worst‑case discounted reward, with the ambiguity set characterized as a Kullback–Leibler ball centered at the posterior beliefs. We derive a dynamic programming relation for the worst‑case discounted reward and introduce a practical backward approximation algorithm to compute the optimal robust policy. Extensive experiments on synthetic and real lending data show that the proposed method consistently improves robustness and lower-tail performance relative to competitive baselines, at only a modest cost in nominal reward.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.