Detecting Harm Before Data Pooling
Abstract
Pooling another center's data can lower a model's accuracy on its own population when the added data carry a distribution shift. We turn will adding this source help? into a decision a center can make before any joint training and without exchanging raw data. The recipient either pools when the donor is expected to improve its deployment performance, declines when the donor is expected to hurt it, or abstains when the available evidence is insufficient to make either decision. To enable this three-way decision, we define the recipient-specific pooling effect as the change in the recipient's deployment performance that would result from adding the donor's data to its training set, and estimate it using a single aggregated gradient returned by the donor at the recipient's model. This allows the recipient to assess the prospective effect of pooling without accessing the donor's raw records or performing joint training. We provide guarantees for detecting harmful pooling and reducing decision costs under stated conditions. Across five datasets spanning synthetic, natural-image, and medical-imaging settings, our method more accurately identifies beneficial and harmful data additions than baseline approaches and yields reliable pool, decline, and abstain decisions across varying levels of distribution shift.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.