Personalized Estimator Aggregation for Off-Policy Evaluation
Abstract
Off-policy evaluation enables accurate assessment of counterfactual policies using only logged interaction data. Although many estimators have been proposed, their estimation accuracy depends strongly on factors such as the logging policy, the evaluation policy, and the available sample size. This sensitivity makes it difficult to deliver an estimator that performs well across diverse practical settings. Although existing approaches to estimator selection or aggregation attempt to address this challenge, they critically overlook contextual information. In practical applications, the most effective estimator can differ substantially across distinct contexts or user segments. We thus propose Personalized estimator aggregation for Off-Policy Evaluation (PerOPE), a novel framework that adaptively aggregates multiple base estimators to deliver the most suitable one for each individual or context. By accounting for contextual variation, this approach captures the heterogeneity of estimator accuracy across context groups and enables personalized aggregation of estimators based on user features. We conduct extensive experiments on synthetic and real-world production datasets, and the results show that PerOPE consistently outperforms existing estimator selection or aggregation methods across a broad range of experiment conditions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.