acceptodds
Under review as a conference paper at ICLR 2027

Improving Shapley Data Valuation for Recommendation Systems by Debiasing Popularity

Abstract

Data valuation decides which interactions a recommendation system keeps and acts on, and how revenue is split among content providers. A wrong value therefore propagates into both what users see next and who gets paid. In systems with implicit feedback, the logging policy makes such errors systematic: exposure is missing not at random (MNAR), so a point can look valuable because it was shown often rather than because it is relevant. We formalize this intuition by revealing that *naive Shapley valuation* shifts credit from low-exposure to high-exposure items, and can reverse their order when exposure gaps outweigh relevance gaps. To counter this, we propose *SHADE (SHApley with Debiased Estimates)*, which corrects each per-point value with debiased and doubly robust estimates before solving the cooperative game with the closed-form KNN-Shapley recursion. Experiments on synthetic MNAR benchmarks and real recommendation datasets confirm that *naive valuation* consistently under-values the long tail, while *SHADE* recovers much of the oracle level and improves value ranking. The correction also assigns value to items that were never exposed, which enables augmentation and collection strategies that naive valuation cannot support.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.